← Blog

Overconfidence and the illusion of control

Overconfidence is not a mood. It is a gap that can be measured — the range a person gives is narrower than their accuracy warrants — and it is hardest to notice where feedback is slow and noisy, which describes investing. Beside it sits the sense that acting on something makes an outcome likelier, and together they explain how more research raises confidence while leaving accuracy where it was.

What overconfidence actually names

Overconfidence, in the work that named it, is a calibration failure rather than a mood. Asked for a range wide enough to be almost certain of containing an answer, people give one that contains it much less often than that. The range is narrower than the accuracy behind it justifies.

You can run the test on yourself in five minutes, and it is worth doing before reading further. Take ten questions you half-know — the year the Bombay Stock Exchange was founded, the length of the Brahmaputra, the number of districts in Maharashtra — and for each one write a low figure and a high figure, spaced so far apart that you would be genuinely startled to find the truth outside them. Then look all ten up.

Marc Alpert and Howard Raiffa reported the result of exactly that design, in a chapter published in Kahneman, Slovic and Tversky's 1982 volume Judgment Under Uncertainty: the true value landed outside the stated ranges far more often than the stated level of certainty allowed. The number of ranges that missed is in the chapter and is not repeated here, because the direction is the durable part and the size varies with the questions asked.

Notice what the failure is not. It is not ignorance about the Brahmaputra — nobody expected the midpoints to be right. It is that the intervals came back at the wrong width, and widening them was free. Nothing stopped anyone writing a range four times as wide, and the reason nobody did is the whole subject of this article.

An interval is a second-order claim: not an estimate of the quantity, but an estimate of how good your estimate is. That distinction does a great deal of work. Someone can know very little about a company and still be well calibrated about it, by giving a wide range and meaning it. Someone can know a great deal and be badly calibrated, by giving a narrow one. Knowledge and calibration are separate axes: a judgement can be strong on one and weak on the other, and the two are scored by different evidence.

Three different things go by the name

Part of why the word is slippery is that it covers three claims which are measured differently and do not move together. Don Moore and Paul Healy set the distinction out in a 2008 paper in Psychological Review, and it is the single most useful thing to take from the literature.

What it is calledThe claim being madeWhat you would need to score it
OverestimationI did better than I actually didYour predicted result against your actual result
OverplacementI am better at this than most people areYour result against everyone else's results
OverprecisionMy range is this narrowHow often the truth landed inside the ranges you gave

Overplacement is the one everybody has heard of, usually through Ola Svenson's 1981 study in Acta Psychologica reporting that drivers rated themselves safer and more skilful than their fellow drivers. It is also the least stable of the three. Moore and Healy describe a pattern in which the direction flips with difficulty: on hard tasks people tend to overestimate their own score while placing themselves below others, and on easy tasks the reverse. A single label was covering two opposite behaviours.

For a portfolio, the third one is the one that costs money, and it is the least discussed. Overprecision is what decides position size. How much of a portfolio a holding is worth depends entirely on how wide the range of outcomes is taken to be, so a range reported too narrow produces a position sized too large — not because the direction was wrong, but because the uncertainty was understated.

That is also why “I was right about the company” and “I sized it sensibly” are separate questions that the same person usually answers as one. Being right about direction is a first-order judgement. Sizing is a claim about your own error bars, and nothing about a correct call tells you anything about how wide those should have been.

Why investing is a poor place to learn calibration

Calibration is learnable. Weather forecasters are the standard demonstration: Allan Murphy and Robert Winkler, writing in 1977 in the Journal of the Royal Statistical Society, reported that professional precipitation forecasts lined up closely with the frequencies that followed. Days on which forecasters said rain was likely were days on which it rained about that often.

Look at what the forecaster's job supplies that a portfolio does not. The prediction is stated in advance, in a fixed form, with a number attached. The outcome arrives tomorrow. It is unambiguous — it rained or it did not. Nobody can revise what was forecast after the fact, because it was published. And the exercise repeats hundreds of times a year, so a systematic tilt shows up as a pattern rather than as an anecdote.

Daniel Kahneman and Gary Klein, in a joint 2009 paper in American Psychologist written specifically to find where their positions agreed, set out the conditions under which skilled intuition develops at all: an environment regular enough to contain valid cues, and enough practice with feedback to learn them. Robin Hogarth had earlier named the two ends of that scale — kind learning environments, where feedback is prompt and accurate, and wicked ones, where it is delayed, noisy or missing.

An investment decision fails every one of the forecaster's conditions at once. The prediction is usually never stated, so there is nothing to score. The horizon is years, so the outcome arrives after the reasoning has faded. The result is confounded with everything else that happened in those years. And the sample is small: genuinely independent decisions arrive at the rate of a handful a year over a working life, where the forecaster's arrive daily.

Which produces the mechanism this article is built around, and we state it as our reading rather than as anyone's finding. Confidence is fed by the reasoning — by how coherent the case felt while you assembled it. Accuracy is fed by feedback. In a wicked environment, only one of those two inputs is being supplied, so the two quantities drift apart with nothing to pull them back.

The specific error worth naming is treating a profitable outcome as a scored prediction. A holding that rose confirms the reasoning only if the reasoning said what would happen, over what period, and would have been marked wrong otherwise. Absent that, the rise is compatible with a good thesis, a bad thesis and no thesis, and it feels like confirmation in all three cases.

More research, more confidence, the same accuracy

If accuracy has no reliable input in this environment, the obvious response is to work harder on the reasoning — read more, model more, wait for one more quarter of numbers. That does something. What it does is not what it feels like.

Stuart Oskamp ran the cleanest version of this in 1965, published in the Journal of Consulting Psychology. His judges — practising clinical psychologists alongside psychology students — were given a case in stages, answering the same questions and rating their confidence after each new instalment of background. Confidence climbed steadily as the file grew. Accuracy did not follow it up.

Crystal Hall, Lynn Ariss and Alexander Todorov reported the same shape in a different setting in 2007, in Organizational Behavior and Human Decision Processes: given more statistics about the teams involved, people predicting basketball outcomes became more confident without becoming more accurate. Paul Slovic's much-cited study of horse-race handicappers given progressively more variables points the same way, and is worth naming with a caveat — it circulates as settled fact but was not published in a peer-reviewed journal.

The mechanism is not that extra information is worthless, and the account that follows is a reading of those results rather than something either paper measured. Extra information is mostly correlated with what you have rather than independent of it. The fifth broker note repeats the first one's framing; the eighth quarter of margin data restates the trend the first six established. Each new item adds almost nothing that is independent, and adds its full weight to how substantial the case feels. Redundancy is invisible from the inside, because a fact that agrees with the others does not announce that it is not new.

This is where the two mechanisms in this article meet a third. The direction the search runs in was set before the search began, and what gets admitted afterwards is supervised — an anchor sets the question and confirmation decides which evidence answers it. Redundant confirming material is precisely what a positive-test search produces, and it is also exactly what an overprecise interval is built from.

The second mechanism: action that cannot reach the outcome

Overconfidence is a claim about the width of a range. The illusion of control is a different thing: the sense that doing something makes an outcome more likely when the action cannot affect it.

Ellen Langer named it in a 1975 paper in the Journal of Personality and Social Psychology. In the demonstration people best remember, office workers bought lottery tickets; some chose their own, some were handed one at random. Asked later what they would sell the ticket for, those who had chosen wanted considerably more. The draw was identical either way. Across her studies the pattern was that cues normally associated with skill — choosing, familiarity, competition, being actively involved — raised the felt control over an outcome decided by chance.

Treat the size of this one with more caution than the rest of the article. A 1996 review by Clark Presson and Victor Benassi, pooling the earlier experiments statistically — a meta-analysis — found the effect was not produced uniformly across designs, and Suzanne Thompson and colleagues argued in Psychological Bulletin in 1998 that what is going on is narrower than the popular reading — a heuristic that reads control off two things, whether you intended the outcome and whether there is any apparent connection between your action and it. The mechanism is documented. The claim that it operates everywhere is not.

Markets are unusually good at supplying the cues Langer identified, because almost everything about the act of investing is skill-shaped. You choose. You research. You time the entry. You watch. You are competing against someone. Every one of those is real work, and none of them is evidence that the work reaches the part that decides the outcome.

The actionWhat it genuinely controlsWhat it cannot reach
Choosing this scheme over that oneWhat you own, its cost, and its mandateWhat the underlying market does next
Reading ten reports before buyingHow well you understand the businessHow accurate your estimate of its future is
Waiting for a better entry levelYour cost of acquisition, and your tax on exitThe path the price takes afterwards
Checking the app every morningHow many separate evaluations you makeThe value on the screen
Placing the order by hand rather than by mandateThe date on the contract noteThe return between now and the goal
Adding filters to a screenHow short the shortlist isWhether the survivors do better than the ones removed

The middle column is not empty, and that is the point of laying it out this way. Costs, allocation, contribution rate, tax treatment and how often you look are all genuinely inside your control, and they are the least glamorous items on any list. What the illusion does is not invent control out of nothing. It spills the real control in the middle column over into the right-hand one, so effort spent on things that respond gets credited to things that do not.

There is a small piece of direct evidence from a trading floor. Mark Fenton-O'Creevy and colleagues reported in 2003, in the Journal of Occupational and Organizational Psychology, that investment-bank traders who scored higher on an illusion-of-control task were rated lower by their managers on performance and risk management. It is a small professional sample rather than a general law, and it is suggestive rather than decisive — but it is one of the few places the two ends have been measured on the same people.

What the two produce together in an account

Put a narrow interval next to a sense that acting improves the odds and the natural output is activity. A narrow range makes the case look clearer than it is; the control cue makes acting on it feel productive; and neither of the two supplies a reason to wait.

The most-cited account-level evidence is Brad Barber and Terrance Odean's 2000 paper in the Journal of Finance, with the title that did most of the work: trading is hazardous to your wealth. Working through a large sample of household brokerage accounts, they found that gross returns across households were far closer together than net returns, and that the households which turned their portfolios over most heavily kept the least. The same authors followed it in 2001 in the Quarterly Journal of Economics with a study comparing turnover between men and women and reaching the same direction.

Separate the measurement from the reading, because they belong to different tiers of claim. That heavier turnover went with lower net returns in that sample is what was measured. That overconfidence caused the turnover is the authors' interpretation, it is the reading usually attached to the result, and it is contested. And it is US data from a particular period; we are not citing an Indian account-level study here, because we do not have one to cite.

The arithmetic underneath the result needs no psychology at all, which is what makes it hard to argue with. Every round trip pays a spread, brokerage, and the taxes and charges that attach to a transaction. Those costs are certain and immediate; the edge that justifies the trade is neither. Repeat that exchange often enough and the certain side compounds against you regardless of how good the individual calls were.

None of which makes activity itself the error, and the mirror mistake is real. A rebalancing rule trades. A goal reached early is worth acting on. The difference is what triggered it: a rule written in advance fires on a stated condition, while an interval that felt narrow this morning fires on how the case feels. Related mechanisms produce the same activity from different directions — a crowd supplies the sense that everyone else already knows, and a position's distance from cost supplies a reason to act that has nothing to do with the asset.

Write the prediction down, because memory edits it

The useful correction here is not an attitude. It is a record, and the reason is specific: the thing calibration needs is a prediction that cannot be revised after the outcome, and memory does not supply one.

Baruch Fischhoff and Ruth Beyth demonstrated the problem in 1975 in Organizational Behavior and Human Performance. People who had assigned probabilities to events were asked afterwards to recall the probabilities they had themselves given. The recalled figures had moved towards what actually happened. Not invented — remembered, sincerely, in a shifted position.

That single finding is why self-assessment from memory cannot work, however honest the person doing it. Reviewing your own past calls without a record is asking a witness who has already read the verdict. The confident calls that failed are remembered as having carried more doubt than they did, so the tally comes back reassuring and nothing is learned. The evidence for correction has been quietly rewritten in transit.

A written prediction is scorable only if it contains four things, and the missing one is almost always the third. What you expect, stated so that someone else could check it. By when. How sure you are, as a number — seven times in ten, or nine. And what observation would mark it wrong. Without the third, everything can be read as having half-happened; without the fourth, nothing can ever resolve.

The evidence that scoring improves anything comes from tournaments rather than from portfolios. Barbara Mellers and colleagues reported in Psychological Science in 2014 that training in probabilistic reasoning, together with regular scoring and feedback, improved accuracy among participants forecasting geopolitical events. Whether that transfers to decisions taken over years, with your own money and no scoring authority, has not been shown, and we are not going to claim it.

What can be claimed is narrower and enough. A log makes calibration observable, which memory cannot. The interesting output is not whether the individual calls were right — it is what fraction of the entries you marked “almost certain” came true, because that is the number that tells you whether your ranges are the right width. And it costs something real: a file of your own errors, in writing, most of it dull, and it takes years before it says anything.

What this does not fix

Three limits, stated flatly, because a section on correcting a bias that promises more than the evidence supports is performing the bias.

Reading this does not remove it. Emily Pronin, Daniel Lin and Lee Ross documented in 2002 what they called the bias blind spot: people rate themselves as less susceptible to standard cognitive biases than others are, including after the bias has been described to them. The realistic gain from an article like this is noticing one specific decision, not immunity from the class.

The magnitudes are contested and the article has deliberately quoted none. The interval studies replicate in the laboratory; the illusion of control does not appear uniformly across designs; the account-level results are interpretations of turnover data from other markets. Treat the mechanisms as the sturdy part and any number attached to them as an estimate from one setting.

And the correction has a cost with a name of its own. Wider intervals mean smaller positions, more “I do not know”, and slower decisions, including the ones where fast and certain would have been right. Under-confidence is not the safe side of the error — it is the same calibration failure pointing the other way, and it is cheaper only because its costs never show up as a transaction.

There is one more trap in the opposite direction, and it follows straight from the middle column of the table above. Concluding that nothing is controllable is as wrong as the illusion, and more expensive: costs, allocation, savings rate and how the money is taxed all respond to action, reliably and permanently. The mechanisms described here put the return on effort in that middle column rather than the right-hand one; they do not say effort is pointless. Where the emotional weather makes that redirection hard is the subject of its own article, and the full survey of the biases in this family sits in the pillar piece.

Numbers that do not care how confident you were

Everything above turns on a range being reported too narrow, which means the practical follow-up is looking at a spread of outcomes rather than a single figure. A distribution is the format that refuses to be narrowed.

FNOTrader's Mutual Funds app runs on the full AMFI NAV history — around 34 million NAV rows — and reports rolling returns across every available start date rather than one: the worst window, the share of windows below zero, and the spread between best and worst. It also simulates SIP and lumpsum on any scheme and period, reporting the internal rate of return for cashflows on irregular dates — XIRR — against invested value, and the deepest peak-to-trough fall along the way.

The worst window is the figure that matters for this article. It is the historical record's answer to the question an overprecise interval skips: not what the scheme returned, but how badly a holding period could have gone. A prediction written against that spread can be scored later; a prediction written against a single average cannot.

None of it is needed to do the thing described here. A dated note with a number on it, written before the position exists, is the whole method. Past performance is a record of what happened, not an indication of what will. FNOTrader is not a SEBI-registered investment adviser and does not give investment advice.

Common questions

What is overconfidence in investing?

It is a calibration failure rather than a personality trait: the range someone gives as almost certainly containing an answer turns out to contain it far less often. Alpert and Raiffa demonstrated it with general-knowledge questions in work published in 1982. The estimate itself may be reasonable — what is wrong is the reported width of the uncertainty around it.

What is the illusion of control?

The sense that acting on something makes an outcome more likely when the action cannot affect it. Ellen Langer named it in 1975, reporting among other results that people who chose their own lottery ticket valued it more highly than people handed one at random. Cues normally associated with skill — choosing, familiarity, competition, involvement — raised the felt control over outcomes decided by chance.

Are overconfidence and the illusion of control the same thing?

No, and they act at different points. Overconfidence is a claim about how wide your uncertainty is; the illusion of control is a claim about whether your action reaches the outcome. They compound because a narrow range makes a case look clear and the control cue makes acting on it feel productive, so neither supplies a reason to wait.

Why does more research not make me more accurate?

The usual reading — a reading of the results rather than a measurement in them — is that additional information is mostly correlated with what you already had, while feeling like new weight in the file. Oskamp reported in 1965 that confidence rose as a case file grew while accuracy did not, and Hall, Ariss and Todorov found the same shape in 2007 with sports predictions. Redundancy is invisible from the inside, since a fact that agrees with the others does not announce that it is not new.

Why is investing such a bad place to learn calibration?

Because it fails every condition that lets a forecaster improve. The prediction is usually never stated, the outcome arrives years later, the result is confounded with everything else that happened, and the sample of genuinely independent decisions in a lifetime is small. Kahneman and Klein set out the conditions for skilled intuition in 2009 — a regular environment and practice with feedback — and an investment decision supplies neither.

Does keeping a prediction log improve returns?

Nobody has shown that, and this article does not claim it. What a log does is make calibration observable, which memory cannot: Fischhoff and Beyth reported in 1975 that recalled probabilities drift towards what actually happened, so a review done from memory is done by a witness who already knows the verdict. The scorable output is the fraction of your 'almost certain' entries that came true.

What makes a prediction scorable?

Four things, and the third is the one usually missing. What you expect, stated so that someone else could check it; by when; how sure you are, as a number; and what observation would mark it wrong. Without a probability everything can be read as having half-happened, and without a falsifier nothing ever resolves.

Is trading less often the answer?

The measured result is narrower than that. Barber and Odean reported in 2000 that among US household accounts, gross returns were much closer together than net returns and the heaviest traders kept the least — costs are certain and immediate while the edge is neither. Reading that as overconfidence is the authors' interpretation and is contested, and a rebalancing rule trades too. What separates the two is whether a stated condition fired or the case simply felt clear.

Does knowing about these biases stop them?

The evidence points the other way. Pronin, Lin and Ross documented in 2002 that people rate themselves as less prone to cognitive biases than others, including after the bias has been explained. That is why the response described here is procedural — writing a prediction down with a number and a date before the position exists — rather than a resolution to be more objective.

Continue reading

More in Behavioural Finance · App: Mutual Funds · Definitions: glossary · Free tools: calculators · All: every article