Human-AI Reliance: When Help Becomes Dependence
Reliance is the degree to which a user lets an AI system's output influence or replace their own judgment, observed in behavior rather than inferred from expressed confidence. This article treats reliance as a behavioral construct, separates appropriate from over- and under-reliance using ground-truth-aware experimental designs, and locates the productivity/skill-atrophy debate inside that frame. The strongest current claim is conditional: high reliance is not pathological by definition, but stable patterns that selectively increase acceptance of wrong outputs, suppress verification, or erode unaided competence in domains where the user retains responsibility do warrant the stronger language of dependence.
Coverage note: verified through May 8, 2026.
Reliance as behavioral delegation
The literature on human-AI interaction has accumulated a thick vocabulary — trust, confidence, acceptance, calibration, dependence, automation bias, algorithm aversion, adoption — and the words are routinely conflated. Reliance is the most behaviorally specific of them, and treating it that way is the first methodological move on which the rest of the topic depends.
Reliance is the observable delegation of judgment. A user relies on an AI system when they let its output substitute for, anchor, or shortcut their own decision process: by accepting an answer, by failing to override a flawed suggestion, by copying generated code without inspection, by selecting from an AI-ranked option set without considering rejected alternatives, by allowing the system to determine the search space of what counts as a candidate solution. None of these acts requires the user to believe anything in particular about the system. A user with low expressed trust may still rely heavily if independent verification is more expensive than accepting the answer. A user with high expressed trust may still verify carefully because their professional accountability does not let them outsource judgment.
This separation matters because attitudinal measures and behavioral measures dissociate routinely. In a preregistered study (N = 404) on uncertainty expression in large language models, Kim, Lee, and colleagues at Microsoft Research found that first-person uncertainty markers in answers to medical questions reduced user agreement with the model and improved overall medical-answer accuracy, while leaving expressed trust largely unchanged in some conditions (Kim et al., 2024). The behavioral effect was the load-bearing one; the attitudinal effect was incidental. Articles that fold reliance into "trust calibration" lose the ability to detect this kind of dissociation, which is precisely where the action is.
The first practical consequence is that the dependent variable in any reliance claim must be a behavior, not a sentiment. "Users trust AI too much" and "users delegate to AI more often than the evidence warrants" sound like the same statement; they are not. The first is an attitudinal claim recoverable by survey; the second is an empirical claim requiring ground truth about both the AI's outputs and what the user did with them.
The second consequence is that delegation is not the same as helplessness. People delegate to compilers, search engines, calculators, type checkers, and institutional memory without the literature treating those reliances as pathological. AI systems differ from those tools in degree more than in kind — they produce fluent, underspecified judgments in domains where correctness is hard to verify and where the system can absorb not only execution but problem framing, search strategy, and error detection. That makes them unusually capable of displacing the very capabilities needed to supervise them. But this is a domain claim, not a general claim about delegation, and the article that confuses those two will moralize too quickly.
Trust Calibration is therefore an adjacent construct, not the construct: it explains why a user might delegate, but the dependent variable in reliance research is the delegation itself.
Appropriate, over-, and under-reliance: the measurement frame
Once reliance is operationalized behaviorally, the canonical taxonomy is straightforward. Cross user acceptance with ground truth about the AI output, and four cells appear:
| AI output correct | AI output incorrect | |
|---|---|---|
| User accepts | Appropriate reliance | Over-reliance |
| User rejects | Under-reliance | Appropriate non-reliance |
This 2×2 is empirically tractable when ground truth is available — adjudicated answers, expert labels, downstream outcomes — and when a human-only baseline can be established for the same items. Without those two ingredients, every interesting claim about reliance reduces to claims about agreement rates, which confound rational delegation, automation bias, and conformity in ways the literature has spent two decades trying to disentangle.
Vasconcelos and colleagues (CSCW 2023) anchor the contemporary behavioral framing. Across five studies (N = 731), they argue that overreliance on AI is not best modeled as naive trust; it is a cost-benefit choice in which users decide whether to engage with explanations and verification given task difficulty, explanation difficulty, and incentives. When verification is cheap or the task is easy, users engage more and overreliance falls. When verification is expensive relative to the perceived stakes, users shortcut by accepting outputs — including wrong ones. Their formal framing matters because it predicts that "make explanations richer" will not, on its own, reduce overreliance; the intervention has to lower the cost of verification, or raise the perceived value of catching an error, or both (Vasconcelos et al., 2023).
This recasts several findings that earlier framings treated as evidence of credulity. Users who accept fluent AI output without checking it may be acting credulously; they may also be acting rationally given a verification cost they have correctly priced. Distinguishing those two requires manipulating the cost structure or the ground-truth distribution, not asking users how confident they were.
The frame also exposes the symmetric failure. Under-reliance is not the morally safe alternative to over-reliance: it is the same construct mismatched in the opposite direction, and it has its own well-documented sources. The 2×2 is symmetric on purpose. An article that treats only the upper-right cell as a problem has imported a value judgment that the measurement frame does not carry.
The frame has known limits. It assumes ground truth exists and is observable; it assumes a usable human baseline; it assumes errors are roughly comparable in cost across cells. None of those holds universally. In safety-critical domains, the cost of accepting an incorrect output can be orders of magnitude larger than the cost of rejecting a correct one, and the 2×2 has to be weighted accordingly. In domains where ground truth is contested — diagnosis, judgment under uncertainty, creative writing — the frame degrades into the same agreement-rate confound it was meant to escape. A reliance claim in those domains has to either smuggle in a normative criterion (expert consensus, downstream outcome, internal consistency) or admit that "reliance" cannot be cleanly measured.
The narrow form of trust calibration belongs here. Behavior is calibrated when delegation tracks the system's actual comparative value under the task's cost structure, not when expressed trust matches an abstract probability of AI accuracy. This narrower definition keeps calibration useful as a measurement anchor while preventing it from absorbing the entire topic.
Experimental designs that separate reliance from confidence
Three families of experimental design recur in the reliance literature, and they are not interchangeable. Knowing which design produced a finding is the difference between reading evidence and reading agreement statistics.
| Design | What the participant does | Primary measure | What it can show | What it confounds |
|---|---|---|---|---|
| Compliance task | Sees AI advice, then decides | Acceptance rate, conditional on AI correctness | Whether users accept correct vs incorrect AI outputs at different rates | Anchoring on AI; cost of verification; expertise floor |
| Override task | Forms an initial judgment, sees AI advice, may revise | Switch rate, switch quality | Whether users override based on evidence or anchor on AI | Order effects; perceived authority of revision |
| Agreement task | Sees AI output (or AI summary) and produces a final answer | Final-answer agreement; textual similarity | Conformity; downstream contamination | Severe — without ground truth and a human baseline, agreement ≠ reliance |
| Consult-or-not | Decides whether to invoke the AI before producing a judgment | Selection rate; downstream accuracy | Strategic delegation | Selection bias by task difficulty; pre-AI confidence |
Compliance designs are the workhorses for measuring over- and under-reliance directly. They are also the easiest to misread: a high acceptance rate is meaningless without conditioning on whether the AI was correct. The headline statistic — "users accepted X% of AI suggestions" — is exactly the kind of finding that supports several mutually incompatible interpretations.
Override designs add a useful constraint by requiring the user to commit before exposure. They isolate the update the AI causes. A user who switches toward correct AI advice and away from incorrect AI advice is exhibiting a reliable Bayesian update; a user who switches uniformly toward AI advice is exhibiting anchoring or compliance pressure. Override designs are how the field empirically separates "AI helped" from "AI dominated."
Agreement designs are weakest. Without ground truth and an unaided baseline, an agreement statistic is a conformity statistic with extra steps. Several recent LLM evaluations conflate the two, in part because LLM outputs are easier to compare textually than behaviorally.
Self-report and behavior also dissociate within designs, not just across them. Kim et al.'s uncertainty-expression study used a within-subjects compliance task and explicitly compared behavioral reliance against expressed trust (Kim et al., 2024). The behavioral effect of first-person uncertainty markers — reduced agreement, improved final-answer accuracy — was robust; the trust effect was smaller and less consistent. A study that had only collected trust scores would have produced a weaker, possibly null, conclusion about an intervention that in fact worked.
The methodological corollary: a reliance result that comes only from self-report should be read as a hypothesis about behavior, not as evidence about it. Lee et al.'s CHI 2025 survey of 319 knowledge workers — which found self-reported reductions in critical-thinking effort associated with higher GenAI confidence (Lee et al., 2025) — is illustrative. The finding is suggestive and worth following up. It is not, in its current form, evidence that critical-thinking capacity has actually declined, because the dependent variable is a self-report about cognitive effort, not a behavioral measure of cognitive performance.
Moderators: what changes reliance
Reliance is not a stable trait. It varies systematically with at least five interacting factors, and population-average claims about "users" or "AI" tend to break the moment the factors are stratified.
AI accuracy and visible errors
The most reliable population-level finding is that visible errors reduce reliance — sometimes more than warranted. Dietvorst, Simmons, and Massey (2015) coined "algorithm aversion" for the effect that participants who saw an algorithm err became reluctant to use it even when the algorithm continued to outperform a human forecaster. The asymmetry was striking: a human forecaster who erred at the same rate did not lose the same degree of trust. The mechanism appears to be that algorithmic errors feel more diagnostic — a single mistake is read as evidence of a defective system rather than as ordinary noise (Dietvorst et al., 2015).
This is the mirror image of automation bias. Both errors exist; they fire under different conditions; and an article that treats only one of them inverts the literature.
Task type: objective vs. subjective
Castelo, Bos, and Lehmann (2019) showed that willingness to use algorithmic advice falls sharply for tasks perceived as subjective (recommending art, judging humor, choosing a date) compared with tasks perceived as objective (forecasting, classification), across four online lab studies and two field studies. Crucially, the moderator is perceived subjectivity, not actual task structure: the same task framed differently produced different reliance (Castelo et al., 2019).
The practical implication is that domain claims about reliance need to specify task type. Findings from medical diagnosis or code completion will not transfer to creative work or interpersonal advice, even if the underlying systems are technically capable in both.
User expertise
Expertise is the most-studied and most-contradictory moderator. The intuitive prior — novices over-rely, experts under-rely — is partially supported, partially inverted, and heavily task-dependent.
Logg, Minson, and Moore (2019) documented "algorithm appreciation": across six experiments, lay participants weighted algorithmic advice more heavily than equivalent human advice. Crucially, experts showed the opposite pattern, weighting their own judgment over the algorithm even when the algorithm was demonstrably more accurate (Logg et al., 2019). Expertise inverts the population effect. This is one reason that "AI augments humans" findings depend so heavily on which humans were sampled.
Novice over-reliance is more visible in domains where the novice cannot verify. Studies of CS1 students using LLM-based code generators (Kazemitabaar et al., 2023) document a range of strategies — full-solution prompting, step-by-step prompting, hybrid approaches, and manual work — with patterns shifting by task difficulty and student fluency. The pure "novices accept everything" picture does not survive contact with the data.
Stakes and reversibility
Reliance falls when stakes rise and rises when stakes are low — but the direction reverses if the user judges that the AI is more reliable than they are at high-stakes tasks. The asymmetry of error costs matters more than the absolute stakes. In a domain where rejecting a correct AI output is cheap (an extra minute of verification) and accepting a wrong one is expensive (a misdiagnosis), rational users should under-rely, and observed under-reliance is appropriate. The mirror holds for low-cost, high-volume triage.
Verification cost
Verification cost is the moderator Vasconcelos et al. center, and it is doing more work in the literature than it is usually credited with. When verification is expensive relative to perceived stakes, even experienced users accept fluent outputs they would have rejected on inspection. When verification is cheap — a glance suffices — overreliance falls. Many interventions that look like attempts to "increase trust" or "improve calibration" are most parsimoniously read as attempts to lower verification cost: better explanations, surfaced uncertainty, side-by-side comparisons with alternatives (Vasconcelos et al., 2023).
The five moderators interact. A novice on a high-stakes objective task with a visibly accurate model and cheap verification produces a different reliance pattern from an expert on a low-stakes subjective task with an opaque model and costly verification. Population-average claims that ignore the stratification routinely report a real effect for one combination as if it were a property of "AI reliance" in general.
Algorithm aversion as the mirror of automation bias
The popular framing of AI reliance — fluent outputs hypnotize gullible users — captures one tail of a two-tailed distribution. The other tail is just as well documented and is at least as costly in some domains.
Automation bias and algorithm aversion are not opposed theories; they are complementary descriptions of behavior that differ along the moderators above. Automation bias dominates when the system is presented as authoritative, the user lacks verification capacity, the task is objective, and stakes are moderate. Algorithm aversion dominates when the user has watched the system err, the task is perceived as subjective, the user has invested expertise, or the user faces accountability that punishes algorithmic mistakes more than equivalent human ones.
The cost of under-reliance is easy to underweight in a literature whose imagery favors over-reliance. Patients who reject correct algorithmic diagnoses, judges who reject calibrated risk scores in favor of demographically biased intuitions, hiring managers who insist on personal interviews despite their predictive validity below structured assessments — all of these are reliance failures with measurable harms, and none of them fits a "help becomes dependence" frame. They fit a "rejection of help that would have outperformed unaided judgment" frame, which is harder to moralize.
The right reading is that reliance has two failure modes, not one, and that the population-level direction of the failure depends on the moderators rather than on a fixed psychology of trust. Articles or interventions that target only the over-reliance tail typically displace failure into the under-reliance tail rather than eliminating it.
Interventions and their second-order failures
Most named interventions in the reliance literature target a specific failure mode. Few test for displacement onto the opposite failure mode. The result is a thin layer of evidence that interventions "work" in the sense of moving the headline metric, layered over a thicker layer of evidence that they trade one failure for another.
| Intervention | Targeted failure | Likely displacement |
|---|---|---|
| Confidence/uncertainty displays | Over-reliance on overconfident outputs | Alert fatigue; learned dismissal; under-reliance on well-calibrated outputs |
| Explanations of model reasoning | Over-reliance on opaque outputs | Persuasive explanations create false confidence; verification theater if explanation cost = 0 |
| Forcing functions (mandatory review) | Failure to verify | Rubber-stamping under throughput pressure; perceived paternalism; tool abandonment |
| Friction (delays, second screens) | Snap acceptance of wrong outputs | Productivity collapse; users find workarounds; net throughput unchanged |
| Adversarial framing ("AI says X — do you agree?") | Anchoring | Conformity pressure inverted; users may now reject correct outputs |
Vasconcelos et al.'s argument that overreliance is a cost-benefit phenomenon predicts most of these displacements. Interventions that raise verification's perceived value (cheaper explanation, salient uncertainty) reduce overreliance; interventions that raise verification's perceived cost (extra clicks, overload of caveats) reduce overall acceptance — including of correct outputs. Net task performance, not net acceptance, is the right outcome metric (Vasconcelos et al., 2023).
The Kim et al. uncertainty-expression result is one of the cleaner cases of an intervention that beats the displacement problem. First-person uncertainty markers in LLM outputs reduced incorrect-answer acceptance more than correct-answer acceptance, raising final medical accuracy (Kim et al., 2024). The intervention worked because it was selective: uncertainty rose when the model was actually less reliable. An intervention that uniformly added uncertainty to all outputs would predictably under-perform, because the user has no signal to discriminate cases that warrant more verification from cases that do not.
The most common failure mode of intervention research is testing only against the targeted error. An intervention that reduces acceptance of wrong outputs should also be tested against acceptance of right outputs, against unaided baseline performance, and ideally against a delayed-transfer condition where the support is removed. The literature has been slow to adopt this standard, in part because the field's incentives reward demonstrating an effect more than ruling out displacement.
This is also where verification theater enters. A nominal review step satisfies a process audit without producing independent judgment: the user clicks through, anchors on the AI output, and signs the form. Verification theater is detectable only if the experiment measures the quality of overrides, not just their frequency. A user who overrides 20% of suggestions but switches uniformly toward whatever was salient is not actually reviewing; they are theatrically reviewing. Designs that seed errors and measure detection are the cleanest way to catch this; without seeded errors, the data look like attentive use.
Productivity, verification, and the skill-retention problem
The empirical center of the contemporary reliance debate is software engineering, because (a) the tools — Copilot, Cursor, Claude Code, and successors — have been adopted at scale, (b) the population is technically literate enough to participate in studies, and (c) ground truth in the form of compiling/passing tests provides a cleaner outcome measure than most knowledge work. The evidence is rich, fast-moving, and partially contradictory, and the contradiction is itself a finding.
Short-run productivity gains in controlled settings
The first wave of evidence reported productivity gains. Peng and colleagues at GitHub ran a controlled experiment on Copilot in which participants implemented an HTTP server in JavaScript, finding completion 55.8% faster with Copilot than without (Peng et al., 2023, arXiv:2302.06590). Cui and colleagues reported field experiments at Microsoft, Accenture, and an anonymous partner with thousands of developers, finding a 26% increase in completed tasks per week among developers given Copilot, with larger effects for less-tenured developers (Cui et al., 2024).
Both results have caveats. Peng et al. used a constrained, well-defined task; Cui et al. controlled for some but not all organizational confounders, and the headline number bundles compliance, task selection, and self-reported activity. Neither study measured what the user retained after the assistance was withdrawn.
Mature codebases and the inversion
The METR study of experienced open-source developers, published in mid-2025, ran a randomized within-subjects experiment in which participants completed real tasks on their own mature repositories, sometimes with AI tools and sometimes without. Participants expected and reported speedups; their measured completion times were 19% slower with AI (METR, 2025). The result was widely cited as evidence that the productivity story had been overstated.
METR's own February 2026 follow-up cautioned against reading the 2025 study as a settled finding. As AI tools have improved and as the population of "AI-using developers" has shifted in skill and selection, later measurements have produced different effect sizes; the 2026 update notes that task selection (which kinds of tasks developers choose to attempt with AI) and tool capability are both moving variables, and that comparable studies a year apart are not measuring the same intervention (METR, 2026).
The honest read is that there is no single "Copilot productivity effect." There is a family of effects that depend on task ecology (greenfield vs brownfield), codebase maturity, developer expertise, tool capability, the time horizon of measurement, and what counts as an outcome. Articles that report a single number — for or against — are over-summarizing.
Mechanism studies: how Copilot changes work
Mechanism studies are doing more useful work than headline productivity studies. They describe what changes when developers use AI assistance, which constrains the space of plausible long-run effects.
Barke, James, and Polikarpova (2023) observed two interaction modes among developers using Copilot: an "acceleration" mode in which developers used the tool to produce known code faster, and an "exploration" mode in which they used it to find unfamiliar APIs or approaches (Barke et al., 2023). The two modes have different cognitive profiles: acceleration substitutes for typing and recall; exploration substitutes for search and decomposition. They predict different long-run effects on different skills.
Kazemitabaar et al. studied novices in CS1 contexts, observing patterns ranging from full-solution prompting to step-by-step prompting to manual work, with frequency varying by task difficulty and student capability (Kazemitabaar et al., 2023). The distribution suggests that "novices over-rely" is too coarse: novice behavior diverges by task structure and by individual fluency.
A Stanford SCALE replication studied "brownfield" programming — modification of existing code — and reported a comprehension-performance gap: developers using Copilot completed tasks at higher rates but did not improve in their understanding of the code they were modifying (SCALE, 2025). This is the cleanest available evidence that short-run task completion and longer-run understanding can diverge under AI assistance. It is one study; it has not been broadly replicated; it is suggestive rather than decisive. Its mechanism is plausible: if AI handles the parts of comprehension that previously required active inference (variable tracing, control-flow modeling), the user gets the answer without forming the model.
Self-reported critical thinking effects
Lee et al.'s CHI 2025 survey (N = 319) reported that knowledge workers who expressed higher confidence in GenAI tools also reported lower critical-thinking effort across their workflows (Lee et al., 2025). The result is widely cited, but it is a self-report on both sides — confidence and effort — and it does not measure cognitive performance. The honest claim it supports is that the felt experience of cognitive effort declines with AI confidence, which is consistent with both "users are appropriately offloading routine reasoning" and "users are coasting on AI outputs they should be checking." Distinguishing the two requires a behavioral measure of cognitive output, which the survey design does not provide.
What the evidence supports, and what it does not
The strongest claims the current evidence supports are:
- AI assistance produces measurable short-run productivity gains in well-defined tasks, especially greenfield tasks of moderate difficulty for less-experienced users. The Peng et al. and Cui et al. results survive the caveats.
- Productivity gains do not transfer cleanly to mature, complex codebases, as the METR 2025 study shows. The 2026 follow-up suggests this is partly a property of evolving tools and selection effects, not a stable inversion.
- The mechanism of assistance matters: tools used in "acceleration" mode have different effects than tools used in "exploration" mode, and tools that complete tasks without fostering comprehension produce a comprehension-performance gap.
- Comprehension and task completion can dissociate, at least in brownfield programming. This is the cleanest empirical foothold for the skill-atrophy concern.
The evidence does not yet support:
- A general claim that AI use causes durable skill loss across populations and domains.
- A claim that the comprehension-performance gap, observed in one study, generalizes to all knowledge work.
- A claim that self-reported critical-thinking reductions reflect actual cognitive decline rather than appropriate offloading.
The status of skill atrophy is therefore: plausible mechanism, suggestive single-study evidence, and not yet the kind of repeated longitudinal finding that would justify the strong dependence narrative. The honest report is that the question is open and worth resolving — and that the cheapest decisive design is probably a within-subjects study comparing answer-first AI assistance, scaffold-first AI assistance, and unaided work, with a delayed unaided transfer test that includes seeded AI errors.
When delegation becomes abdication
Up to this point the article has resisted moralizing. The remainder of the topic is where the moral question actually lives, and it lives there because some forms of delegation predictably degrade things that are not reducible to task accuracy.
A tiered claim is more defensible than a single one.
Reliance is morally neutral as an isolated act. Accepting AI help when the system is accurate, verification is expensive, and stakes are low is not a moral failure; it is efficient delegation, indistinguishable from a hundred other tools the user already relies on without anxiety. Treating every act of AI delegation as morally suspect imports a hostility to delegation in general that the rest of the user's working life does not share.
Reliance becomes prudentially risky when it predictably reduces capacities the user has independent reason to maintain. If a developer's reliance on AI tooling reduces their ability to debug under outage, to onboard to a new codebase that the AI has not been trained on, or to recognize when the system is operating outside its competence, the delegation has costs that do not show up in the same-session productivity measure. These are recoverability costs — the user is fine until they aren't — and they are exactly the kind of cost that short-run experiments cannot detect.
Reliance becomes pathological when a stable pattern degrades agency, competence, error detection, or accountability in a domain where the user retains responsibility. This is the strongest moral language the evidence supports, and it is conditional on three things: the pattern is stable, the degradation is downstream and demonstrable, and the user (or the user's role) bears responsibilities that delegation cannot transfer. A physician, engineer, judge, or manager cannot fully transfer accountability to an AI system, regardless of the efficiency of the delegation. A reliance pattern that systematically displaces the cognitive work needed to discharge that accountability is pathological in the specific sense that it is incompatible with the role's constitutive demands — not in some general sense of "people are becoming dependent."
The distinction between efficient specialization and constitutive abandonment does most of the moral work. We delegate execution to compilers without abandoning what programming requires of us. We delegate retrieval to search without abandoning literacy. The question for AI systems is which delegations are like compiler-use (extending capability without replacing it) and which are like the kind of dependence in which the user can no longer recognize when the tool is failing them. The former is morally neutral; the latter is what the word "dependence" should be reserved for.
The honest position is that we do not yet know, for most AI-assisted workflows, which kind of delegation is in play. The mechanism studies above suggest that the answer varies by tool, task, and user, and that the field has not yet measured the long-horizon variables that would distinguish the cases.
Detection signals: when reliance has crossed the line
A reliance pattern is materializing as a problem when one or more of the following are observable:
- Wrong-answer acceptance exceeds the unaided baseline. Users accept incorrect AI outputs at a higher rate than they would have produced the same incorrect answer themselves. This is the cleanest signal of automation bias.
- Override quality is uniform with respect to AI correctness. Users override at some rate, but the rate is the same whether the AI was right or wrong. This is verification theater.
- Acceptance increases without accuracy gains. Users delegate more tasks to AI over time, but task quality is flat or declining. The delegation is increasing without earning the trust it implies.
- Unaided performance declines on matched tasks. Users' performance without AI on a task class drops over a measurement period, controlling for task selection. This is the skill-atrophy signal.
- Time-to-debug rises on AI-assisted code. Even if assisted productivity rises, the time required to fix problems in AI-generated artifacts climbs, suggesting that comprehension has not kept pace with output.
- Responsibility language shifts. Post-hoc explanations move from "I judged this acceptable" to "the AI said it." This is responsibility diffusion, and it is one of the few signals visible in qualitative data when the quantitative metrics are stable.
- Competence-boundary detection fails. Users no longer notice when the AI is operating outside its competence — when the task is sufficiently novel, the data sufficiently out-of-distribution, or the system sufficiently mistaken — and act on outputs they should have flagged.
These signals are not sufficient individually. None of them identifies pathological reliance on its own; each has benign explanations. They are useful in combination, and they are useful as a checklist for what an evaluation framework should actually measure if it claims to be tracking dependence rather than tracking reliance.
The detection problem at the article level mirrors the detection problem at the user level: the failure modes that matter are not the ones visible in same-session data. An honest evaluation framework needs delayed transfer, seeded errors, accountability tracking, and longitudinal measurement. Most current evaluation frameworks have none of these.
Open questions and what would resolve them
Several questions in this topic are currently open in ways that the literature has not yet decisively closed.
Does AI assistance cause durable skill loss, after controlling for task selection and exposure length? The mechanism is credible; the comprehension-performance gap is observed in at least one study; the claim has not been replicated longitudinally in field conditions. The cheapest decisive design is the within-subjects answer-first vs. scaffold-first vs. unaided comparison with a delayed transfer test, ideally pre-registered and conducted across multiple task classes.
Are interventions that reduce over-reliance net positive once displacement is measured? The literature has documented many interventions that reduce wrong-answer acceptance. Few have measured whether right-answer acceptance also fell, whether net task performance rose, and whether users found workarounds. A second-generation intervention literature that systematically measures displacement is overdue.
Does reliance behavior generalize across task types, or are findings domain-specific? Castelo's task-subjectivity result and Logg's expertise-inversion result both suggest the latter, but the literature still routinely cites cross-task generalizations. A meta-analysis stratified by task type, expertise, and AI accuracy would be useful.
At what point does reliance become a property of the workflow rather than the user? Organizational adoption of AI tooling creates lock-in: once a workflow assumes the tool, the marginal cost of working without it rises sharply, and individual reliance is no longer a free behavioral choice. The literature has been slow to take seriously the workflow-level analogue of dependence, in which the unit of analysis is not the user but the team or the institution.
How should accountability be allocated when AI assistance is integral to the decision? This is a question for legal, professional, and ethical frameworks rather than for empirical research, but it shapes what counts as pathological reliance. A profession that allows full transfer of judgment to AI assistance has a different threshold for pathology than one that does not.
Where this leaves the topic
The most defensible synthesis of the current literature is conditional and behavioral. Reliance is observable delegation, measured against ground truth and a human baseline, and judged against a cost structure. Over- and under-reliance are symmetric failure modes; neither is the default. Algorithm aversion is as well-documented as automation bias; the mix depends on moderators that the population-average framing obscures. Interventions that target one failure typically displace it into another, and net task performance — not net acceptance — is the right outcome measure. Productivity gains in software engineering are real in some settings and have not generalized cleanly to others; mechanism studies suggest a comprehension-performance gap that is the most plausible foothold for the skill-atrophy concern, but the longitudinal evidence is not yet decisive.
The moral verdict is also conditional. Reliance is not pathological by default. It becomes prudentially risky when it erodes recoverability, and pathological when it stably degrades agency, competence, error detection, or accountability in a role that constitutively requires them. The line between efficient specialization and constitutive abandonment is the line the topic actually rests on, and the evidence base has not yet measured it cleanly enough to draw that line in most domains.
The article most worth writing on this topic is therefore not a morality tale about how help becomes dependence. It is a measurement primer that earns the moral language only where the evidence supports it, and that holds open the empirical questions whose answers will determine how much of the moral language is justified.
Companion entries
Core theory:
- Trust Calibration
- Automation Bias
- Algorithm Aversion
- Human-AI Decision Making
- Behavioral Delegation
Measurement and methodology:
- Compliance Override and Agreement Tasks
- Ground Truth in Human-AI Studies
- Self-Report Behavior Dissociation
- Within-Subjects Transfer Designs
Practice:
- Verification Theater
- Uncertainty Expression in LLM Outputs
- Explanations and Cognitive Forcing Functions
- Copilot and Coding Assistant Evaluation
- Brownfield vs Greenfield Tool Evaluation
Counterarguments and adjacent constructs:
- Algorithm Appreciation
- Task Subjectivity Effects
- Expertise Inversions in Advice Taking
- Productivity Measurement Under Tool Adoption
Normative:
- Delegation vs Abdication
- Constitutive Practice and AI Substitution
- Responsibility Diffusion
- Recoverability as a Reliance Criterion
Counterarguments to the dependence frame:
- Efficient Specialization vs Skill Loss
- Tool-Use Continuity Arguments
- Under-Reliance as the Symmetric Failure