← Reference · Nestor G Pestelos Jr · Print this page
Behavioral Economics · Intertemporal Choice · Decision Science
Hyperbolic Discounting
Reference entry · last updated 20260924
Hyperbolic discounting is a pattern of intertemporal choice in which the subjective value of a delayed reward falls steeply over short delays and gently over long ones, so the implied discount rate declines as the delay grows. It is contrasted with exponential discounting, in which one constant rate applies at every horizon.[1, 2, 5]
First principles and definitions
Delay discounting is the decline in the subjective value of an outcome as the delay to its delivery increases. The economics literature treats the discounted utility (DU) model, proposed by Paul Samuelson in 1937, as the standard benchmark. Under DU, the value of a reward of amount \(A\) delayed by \(d\) periods falls by a constant proportion per period, giving an exponential function with discount rate \(k\):[3, 4, 6]
\[ V_d = A e^{-kd} \]
An exponential discounter is time consistent: preferences over two future dates do not depend on when the question is asked. Robert Strotz showed in 1956 that a person whose discounting is not exponential will want to constrain their own later choices, which makes commitment mechanisms potentially important determinants of economic outcomes.[5, 7]
Hyperbolic discounting names the family of alternatives in which the discount rate declines with the length of the delay. The most common descriptive form, used in animal and human delay-discounting research, is the hyperbola associated with George Ainslie and James Mazur:[1, 2, 3]
\[ V_d = \frac{A}{1 + kd} \]
Here \(k\) is the rate of delay discounting. The distinction from the exponential case is that the rate is not constant across horizons. Laibson describes a hyperbolic discount function as one with a relatively high discount rate over short horizons and a relatively low rate over long horizons.[5]
Time inconsistency and preference reversal follow from that shape. When both a smaller sooner reward and a larger later reward are distant, the larger one is preferred. As the smaller reward becomes available sooner, its subjective value can rise above the larger one, and the choice flips without any new information arriving. Ainslie's 1975 review argued that the hyperbolic curves fitted to animal choice data predict exactly this reliable change of preference as a function of elapsed time.[1, 5]
The quasi-hyperbolic or beta-delta form, used by Laibson in his 1997 consumption-saving model, describes the same intuition in a computationally simpler way: a present-bias factor downweights all future rewards relative to the present, while a constant factor governs discounting between any two future periods. Hyperbolic and quasi-hyperbolic models make similar behavioral predictions, but their mathematical implications differ, so the literature distinguishes them even when a statement applies to both.[5, 8, 9]
Illustrative example. Under the exponential benchmark, with \(A = 100\) units and \(k = 0.05\), the discounted value of a reward delayed by ten months is 60 units, so a person with that rate is indifferent between 100 units in ten months and 60 units now. At \(k = 0.20\), the same delayed reward is worth 13 units, and the smaller sooner option wins. Those two figures are the ones Madden and Johnson use to illustrate the exponential function.[3] The same rates applied to the hyperbola give 66.7 units and 33.3 units, since the two functions agree closely at very short delays and separate as the delay grows. The figures show how a fitted rate maps onto choice; they are not estimates of any population's rate.
Both the hyperbolic and the quasi-hyperbolic models are descriptive, or positive, accounts of behavior rather than claims about psychological or neural mechanism. Time discounting should also be separated from time preference, since a lower weight on a future consequence can come from uncertainty, changing tastes, or anticipation rather than from impatience alone.[6, 8]
Evidence and limits
Systematic departures from exponential discounting have been reported in human subjects, including people with substance-use disorders, with real and hypothetical rewards and for both delayed gains and delayed losses, and in animal experiments with rats and pigeons choosing between immediate and delayed food. Fitted to such data, the exponential function tends to overestimate discounted values at short delays and underestimate them at long delays, while the hyperbolic function fits more closely.[2, 3]
The exact shape remains an open question. Green and Myerson showed that adding an exponent to the denominator of the hyperbola improves fits for human data, but once that exponent departs from one, the fitted \(k\) can no longer be read as a simple rate of discounting; several rat and pigeon experiments found no improvement from the extra parameter. A hypothesis-free alternative is to compute the area under the obtained indifference points rather than fit a curve at all.[3, 10, 11]
Declining impatience itself has been challenged. Read's three choice experiments found strong evidence of subadditive discounting, meaning that discounting over a delay is greater when the delay is divided into subintervals than when it is left undivided, but no evidence of declining impatience. Read concluded by questioning whether hyperbolic discounting is a plausible account of time preference.[10]
Estimated discount rates also vary enormously across and within studies. Frederick, Loewenstein, and O'Donoghue attribute part of that variation to differences in elicitation procedures and part to the assumption that a single rate can represent considerations that differ across choices, and they record that virtually every assumption of the DU model has been found descriptively invalid in at least some situations.[6]
Meta-analytic work narrows one part of the question. Cheung, Tymula, and Wang pooled 89 papers and 109 estimates of the present-bias parameter. After correcting for selective reporting, the estimated present bias for monetary rewards is close to absent, at 0.98 with a 95 percent confidence interval of [0.979, 0.981], while non-monetary rewards show a much stronger present bias at 0.68 with a confidence interval of [0.57, 0.82]. The authors report substantial heterogeneity across studies arising from elicitation methodology, study location, and reward type, and note earlier findings of no present bias for money once transaction costs and trust in the experimenter are controlled.[9]
Neural evidence does not settle the model. Imaging studies have supported a separate neural systems account, in which one system values immediate rewards and another values delayed rewards. Kable and Glimcher reported behavioral and functional imaging results inconsistent with both the standard behavioral models and that dual-system hypothesis: their participants did not necessarily make the preference reversals predicted by hyperbolic-like discounting, and activity in ventral striatum, medial prefrontal, and posterior cingulate cortex correlated with the subjective value of both immediate and delayed rewards rather than tracking the presence of an immediate option. They propose instead that subjective value declines hyperbolically relative to the soonest currently available reward.[8]
Commitment devices and applications
If preferences are time inconsistent, a person who anticipates the inconsistency has a reason to restrict their own future options. Strotz framed the point in 1956 and listed everyday evidence, including arranged savings plans and the person who "is commonly precommitting" when a garnished income makes saving involuntary. Ainslie's later work developed the same logic as a conflict among successive selves.[1, 5, 7]
Laibson's 1997 model formalizes an imperfect commitment technology: an illiquid asset whose sale must be initiated a period before the proceeds arrive. Illiquid assets behave like the goose that laid golden eggs, since they promise substantial long-run benefits that cannot be realized immediately without a capital loss. The model predicts that consumption tracks income and that consumers show asset-specific marginal propensities to consume. Its welfare implication, that financial innovation raising liquidity can reduce welfare by eliminating commitment opportunities, is a model result rather than a settled empirical finding.[5]
Field evidence links measured discounting to real commitment decisions. Ashraf, Karlan, and Yin designed a commitment savings account for a rural bank in Mindanao and tested it with a randomized design. Of 1,777 surveyed clients, 710 were offered the account and 202, or 28.4 percent, opened one. Women who had shown a lower discount rate on future relative to current tradeoffs in the baseline survey were significantly more likely to open it. After twelve months, average savings balances in the treatment group were 411 pesos higher than in the control group, an 81 percentage point increase over pre-intervention levels. The account restricted access to deposits at the client's own instruction and paid nothing for the restriction, and the authors describe it as designed for people who want to commit and are sophisticated enough to do so.[12]
Policy proposals that follow from present bias, including commitment savings products and default enrollment in retirement plans, depend on the same commitment logic. Their justification is weaker where the meta-analytic evidence finds little present bias for monetary rewards, and stronger where the reward is non-monetary.[9]
Discounting in reinforcement learning
Reinforcement learning defines a discount factor as part of the Markov decision process, and the usual exponential scheme is what yields convergence guarantees for the Bellman equation. Fedus and colleagues note that this convention sits awkwardly against the psychological, economic, and neuroscientific evidence for hyperbolic time preferences, and implement an agent that acts via hyperbolic discounting while still using temporal-difference learning to approximate the hyperbolic discount functions. Separately from the discounting question, they report that learning value functions over multiple time horizons at once works as an effective auxiliary task and often improves on a strong value-based baseline.[13]
This is a design application, not evidence about human choice. Discount functions in an agent are chosen for convergence properties and task performance, and the parallel with human discounting is a modelling analogy.
See also
References
- George Ainslie. "Specious Reward: A Behavioral Theory of Impulsiveness and Impulse Control." Psychological Bulletin 82(4), 463–496, 1975. DOI: 10.1037/h0076860. Free full text.
- James E. Mazur. "An adjusting procedure for studying delayed reinforcement." In M. L. Commons, J. E. Mazur, J. A. Nevin, and H. Rachlin (eds.), Quantitative Analysis of Behavior: Vol. 5. The Effect of Delay and of Intervening Events on Reinforcement Value, 55–73. Hillsdale, NJ: Erlbaum, 1987. Print source; cited here through Madden and Johnson and through Kable and Glimcher, who attribute the hyperbolic form to it.
- Gregory J. Madden and Patrick S. Johnson. "A Delay-Discounting Primer." In Impulsivity: The Behavioral and Neurological Science of Discounting, 11–37. Washington, DC: American Psychological Association, 2010. Free full text.
- Paul A. Samuelson. "A Note on the Measurement of Utility." Review of Economic Studies 4(2), 155–161, 1937. Print source for the discounted utility model; cited here through Frederick, Loewenstein, and O'Donoghue and through Kable and Glimcher.
- David Laibson. "Golden Eggs and Hyperbolic Discounting." Quarterly Journal of Economics 112(2), 443–477, 1997. DOI: 10.1162/003355397555253.
- Shane Frederick, George Loewenstein, and Ted O'Donoghue. "Time Discounting and Time Preference: A Critical Review." Journal of Economic Literature 40(2), 351–401, 2002. DOI: 10.1257/002205102320161311. Free full text.
- Robert H. Strotz. "Myopia and Inconsistency in Dynamic Utility Maximization." Review of Economic Studies 23(3), 165–180, 1956. Print source; quoted here through Laibson's 1997 discussion of precommitment.
- Joseph W. Kable and Paul W. Glimcher. "An 'As Soon As Possible' Effect in Human Intertemporal Decision Making: Behavioral Evidence and Neural Mechanisms." Journal of Neurophysiology 103(5), 2513–2531, 2010. DOI: 10.1152/jn.00177.2009. Free full text.
- Stephen L. Cheung, Agnieszka Tymula, and Xueting Wang. "A Meta-Analysis of Quasi-Hyperbolic Discounting." Working paper, December 2023. Free full text.
- Daniel Read. "Is Time-Discounting Hyperbolic or Subadditive?" Journal of Risk and Uncertainty 23(1), 5–32, 2001. DOI: 10.1023/A:1011198414683. Abstract; full text by subscription.
- Joel Myerson, Leonard Green, and Missaka Warusawitharana. "Area Under the Curve as a Measure of Discounting." Journal of the Experimental Analysis of Behavior 76(2), 235–243, 2001. Print source; cited here through Madden and Johnson, who recommend the measure.
- Nava Ashraf, Dean S. Karlan, and Wesley Yin. "Tying Odysseus to the Mast: Evidence from a Commitment Savings Product in the Philippines." Yale University Economic Growth Center Discussion Paper 917, 2005. Free full text.
- William Fedus, Carles Gelada, Yoshua Bengio, Marc G. Bellemare, and Hugo Larochelle. "Hyperbolic Discounting and Learning over Multiple Horizons." arXiv:1902.06865, 2019. DOI: 10.48550/arXiv.1902.06865.