An investment committee reviews one hundred ideas a year and approves twenty. Of those twenty, twelve deliver returns. The committee reports a 60% hit rate. Not bad.
But is it good? You cannot answer that question with the information given. Suppose the pipeline contained eighty genuinely attractive ideas and twenty poor ones. The committee approved twelve that delivered and eight that did not — but also rejected sixty-eight that would have delivered. Its 60% hit rate conceals a significant miss rate. Alternatively, suppose only fifteen of the hundred ideas were genuinely attractive, and the committee found twelve of them. Now the same 60% hit rate reflects considerable skill. In practice, rejected ideas vanish — you rarely learn whether they would have succeeded. But shadow portfolios and post-hoc tracking of passed opportunities suggest the problem is real, even when the exact numbers are unknowable.
The hit rate isn’t a measure of skill. It’s a tangled composite of two completely different things: how well the committee can distinguish good ideas from bad ones, and how willing it is to say yes. These aren’t the same thing.
What kind of problem is this?
When you are trying to detect something real against a background of noise, and your detection system is imperfect, that’s a classification problem — not a forecasting problem or an optimisation problem. Is this thing I’m looking at real, or is it noise? And the discipline that has formalised the tradeoffs, the errors, and the hidden role of base rates is signal detection theory.
So let’s borrow.
This is the multi-model move: recognise the shape of a problem, find the discipline that has thought rigorously about that shape, and import its frameworks deliberately rather than reinventing them from scratch.
Signal detection theory (SDT) emerged from a collaboration between psychophysics and electrical engineering in the 1950s, codified in Green and Swets’ foundational 1966 text. The original context was literal: radar operators in the Second World War, staring at screens, trying to decide whether each blip was an enemy aircraft or atmospheric noise. Real signals and random noise produced overlapping patterns, and no amount of training could eliminate the ambiguity.
Sensitivity and Criterion
SDT formalises this by imagining two overlapping probability distributions — one for noise alone, one for signal-plus-noise. The overlap is the entire problem. The framework’s central insight is that detection performance decomposes into two independent dimensions.
The first is sensitivity — how far apart the two distributions are. In the technical language, this is d-prime (d’), the standardised distance between the means of the noise and signal distributions. High sensitivity means the world looks genuinely different when a signal is present versus when it isn’t. Low sensitivity means the two states are nearly indistinguishable. Sensitivity is a property of the detection system. You improve it by getting better information, better models, better analytical tools.
The second is the criterion — the threshold at which the observer decides to say “signal.” This is a choice, not a measurement. Slide the criterion leftward (become more liberal) and you catch more genuine signals but also flag more noise. Slide it rightward (become more conservative) and you reduce false alarms but miss more genuine signals. At any given level of sensitivity, this tradeoff is inescapable. The only way to get more hits without more false alarms is to improve sensitivity itself.
Every detection decision produces one of four outcomes: a hit (signal present, correctly identified), a miss (signal present, overlooked), a false alarm (noise mistaken for signal), or a correct rejection (noise correctly dismissed). Note that SDT’s “hit rate” — the percentage of genuine signals you caught — is not the same as the usage from our opening, where “hit rate” meant the percentage of approvals that worked out. These are different numbers, and confusing them is itself a diagnostic error. Plot the true hit rate against the false alarm rate as you sweep the criterion, and you trace the ROC curve — the Receiver Operating Characteristic. A system with high sensitivity bows toward the upper left; one with zero sensitivity produces a diagonal line. The area under the curve is a threshold-invariant measure of discrimination ability, untangled from any particular criterion setting.
This decomposition is SDT’s deepest contribution. In most domains, performance is evaluated as a single number — accuracy, hit rate, the percentage of investments that worked. But a single number conflates two things that should be assessed separately. A conservative investor who approves few ideas and gets a high percentage right might be cautious, not skilled. A liberal investor who approves many and gets a lower percentage right might have genuine sensitivity combined with a willingness to act. You can’t distinguish these cases without pulling the two dimensions apart.
The Base Rate Trap
Suppose genuine investment opportunities — ones that would generate meaningful risk-adjusted returns — constitute 5% of the ideas crossing an investment committee’s desk. The committee has good detection: when a genuinely good idea appears, they recognise it 80% of the time, and they correctly reject 90% of bad ideas. Strong numbers. Skilled committee.
Now do the arithmetic. Out of 1,000 ideas, 50 are genuinely good. The committee identifies 40 of them. Of the 950 bad ideas, it incorrectly approves 95. Total approvals: 135, of which 40 are good. The positive predictive value — the probability that an approved idea is actually good — is roughly 30%.
A 30% success rate from a highly skilled committee. Not because the committee is bad, but because the base rate is low. When genuine signals are rare, even excellent detectors generate mostly false positives. This isn’t an argument against trying. It’s an argument for understanding what the numbers actually mean.
The Payoff Matrix
In The Question Shapes The Answer, we explored Adami’s framework for information: edge requires specifying a target, an observer, and a baseline. SDT adds a complementary layer. Even when you have genuine conditional information — even when your sensitivity is real — the decision to act on that information involves a separate, strategic choice about where to set your threshold. And that choice should depend on the payoff structure and the base rate of genuine opportunity, not on the signal itself.
Consider the asymmetry. For a long-term investor with a diversified portfolio, the cost structure might look like this: missing a genuinely great investment is expensive because opportunity cost compounds over decades, while an underperforming investment drags returns but doesn’t threaten long-term viability. This cost structure argues for a liberal criterion — approve more, accept that many will underperform, because the cost of missing strong investments exceeds the cost of including weaker ones.
For a concentrated, leveraged fund, the cost structure inverts. A false alarm — a position that looked like a signal but was noise — can be catastrophic. The ergodicity problem from The Path The Maths Misses applies: a single bad position can destroy the ability to compound. This argues for a conservative criterion — demand compelling evidence before acting, accept that you will miss many genuine opportunities, because survival dominates.
Both are rational. Neither is “better.” They reflect different payoff matrices applied to the same underlying detection problem. SDT suggests that framing high-conviction and diversified approaches as inherently superior or inferior is a category error. The question isn’t how much conviction to have. It’s what the costs of your errors are, and whether your criterion is calibrated to those costs.
The Invisible Threshold
This connects to something the evolutionary lens from Why Good Strategies Stop Working revealed: investment strategies carry assumptions about their environment. SDT sharpens this. Every organisation carries criterion — a collective threshold for what counts as a sufficient signal to act. It is often implicit, embedded in how many approvals the committee gives, how much evidence is required, what the burden of proof feels like.
Some organisations are miss-averse and set liberal criteria. Others are false-alarm-averse and set conservative criteria. These aren’t personality quirks. They are strategic positions in the sensitivity-criterion space. But they are often invisible — embedded in norms rather than articulated as choices.
The only way to improve both simultaneously is to improve sensitivity — to actually get better at distinguishing signal from noise. If you decompose performance into sensitivity and criterion, you can ask the right question: is our problem that we are poorly calibrated (criterion in the wrong place), or that we genuinely cannot tell signal from noise (low sensitivity)? These require completely different interventions. Recalibrating a criterion is a decision. Improving sensitivity is a capability-building exercise — better data, better models, better analytical infrastructure, the kind of investment in measurement apparatus that The Question Shapes The Answer described.
Post-mortems often confuse the two. When an investment fails, the typical response is to raise the bar — demand more evidence, add another approval layer. This is a criterion shift. If the problem was low sensitivity — if the system genuinely could not distinguish signal from noise — then raising the bar simply means you will miss more signals while still getting fooled by noise that exceeds your new threshold. You have made yourself more cautious without making yourself more skilled. The ROC curve has not moved; you have merely walked along it.
The Detection Question
The discipline SDT offers is in the decomposition. Before asking “should we approve more or fewer ideas,” ask: can we actually tell the difference? Before tightening standards after a loss, ask: was the failure a criterion problem or a sensitivity problem? Before celebrating a high hit rate, ask: what was the base rate, and how many signals did we miss?
In markets, this decomposition interacts with competition. When everyone is trying to detect the same signals, the base rate of genuinely private information drops — more observers means more information gets priced faster. The more efficient the collective detection system, the lower the base rate for any individual. Improving aggregate sensitivity reduces each participant’s base rate, which means even skilled detectors will find that most of their positive identifications are false alarms. Understanding signal detection theory doesn’t make you immune to its dynamics. You will still face ambiguous signals, still set your criterion imperfectly, still be surprised by how often confident calls turn out to be noise.
The decomposition applies wherever imperfect detectors face ambiguous signals. In medicine, a screening test with excellent sensitivity can still produce mostly false positives when the disease is rare — and the response of ordering more tests is a criterion shift, not a sensitivity improvement. In security, tightening protocols after an incident can flood the system with false alarms while doing nothing to improve the ability to detect genuine threats. Any domain where someone is trying to sort signal from noise — and where the cost of errors is asymmetric — is a domain where SDT’s decomposition reveals something that a single accuracy number conceals.
But the framework changes the questions you ask. Not “was I right?” but “can I actually distinguish signal from noise in this domain?” Not “how many decisions worked?” but “what was the base rate, and does my sensitivity exceed chance?”
The discipline is in asking: is this a problem of calibration or a problem of capability?
References & Further Reading
Signal Detection Theory and Psychophysics: David M. Green and John A. Swets
The Signal and the Noise: The Art and Science of Prediction: Nate Silver
Noise: A Flaw in Human Judgment: Daniel Kahneman, Olivier Sibony, and Cass R. Sunstein
Thinking, Fast and Slow: Daniel Kahneman
The Base Rate Book: Integrating the Past to Better Anticipate the Future: Michael J. Mauboussin
What is Information?: Christoph Adami


