We make hundreds of estimates every day. We guess how fast a car is approaching before pulling out into traffic, weigh the invisible toll a risky decision might take on our reputation, or try to balance competing priorities when choosing a career path. But for important, complex decisions, there is a massive benefit between "feeling" an answer and numerically expressing priorities in order to take action. Welcome to the challenge of quantifying judgment. We estimate the likelihood of events happening.
Whether we are dealing with objective physical realities or deeply subjective human values, translating human intuition into data is notoriously difficult. How do we measure the intangible? How do we standardize a gut feeling? How do we communicate our priorities to other decision stakeholders, and compare our areas of agreement and disagreement?
The scientific and most effective way to quantify judgment in order to make decisions under conditions of uncertainty is to base decisions on the probability distribution of what is likely to happen. However, there is no data for the future! Using human judgment to estimate probability distributions for future events applies to any decision under conditions of uncertainty — from evaluating risk exposure to deciding what stock options are a good investment.
The following exercise will illustrate how this is done.
Humans are notoriously bad at absolute measurements, but we're world-class at relative comparison.
If you ask someone, "How many grams does this book weigh?" they'll guess wildly. But if you put a book in their left hand and a slightly heavier book in their right hand, they can immediately tell you, "The left one is lighter."
Pairwise comparisons exploit this biological cheat code. By asking, "Between Project A and Project B, which is more critical right now?" you bypass the cognitive friction of absolute scoring and tap directly into intuitive human judgment.
When rating items individually, it's easy to be lazy. You can comfortably label five different features as "High Priority" or give them all 4 out of 5 stars.
Pairwise comparisons remove this escape hatch. Because every matchup is a forced choice, you can't say everything is equally important. It acts like a tournament bracket for your priorities: only one item can win each match unless the judgments are that they are equal. By the end, definitive, mathematically sound estimates of priorities or likelihoods emerge naturally.
How do you measure things that don't have physical units, like brand reputation, culture fit, or cyber risk severity?
Because pairwise comparisons rely on relative preference rather than absolute values, they allow you to convert purely qualitative gut feelings into precise, numerical priorities or likelihoods. By answering a series of simple "Which is more ___ than the other and by how much" questions, the underlying math used in the Analytic Hierarchy Process — transforms subjective judgment into defensible, priorities or likelihoods.
The best way to feel how well this works is to perform the exercise yourself. This exercise lets you experience firsthand how simple pairwise comparisons using words can accurately estimate priorities and likelihoods.
You'll be shown five shapes — a circle, triangle, square, diamond, and rectangle — and asked to estimate the relative sizes (areas) with pairwise comparisons (two at a time) using words. You'll then be able to compare your estimates with the known relative sizes of the shapes. Think of the area of each shape as the relative importance of an objective when making a decision, or the relative likelihood of the price of a stock at some future time.
This experiment and real-world decisions alike have been performed thousands of times, and can be done individually or with a group.
Try the Area Validation Exercise by clicking below:
See How it Works!
After doing the exercise, see if your results are similar to those below
After doing the exercise above, your results should be similar to those of a group of executives at the Ford Motor Company, shown below:
|
Shape |
Rank |
Proportion |
Pairwise Verbal |
Actual |
|
Circle |
1 |
33.3% |
49.6% |
47.5% |
|
Triangle |
5 |
6.7% |
4.8% |
4.9% |
|
Square |
2 |
26.7% |
23.6% |
23.2% |
|
Diamond |
3 |
13.3% |
14.5% |
15.1% |
|
Rectangle |
4 |
20.0% |
7.5% |
9.3% |
Note how close the estimates are in the pairwise verbal column (column 4) to the actual column (column 5). Note also how deficient the estimates are if one were to simply derive the estimates based on the ranking (ordinal measures) of the shapes (column 3 above).
The accuracy of the derived ratio-scale priorities is truly amazing considering that the inputs were ordinal measures (words on the fundamental AHP verbal scale). Deriving ratio-scale measures from ordinal inputs is somewhat magical since ratio measures have all of the information of ordinal measures, plus interval and ratio meaning as well. In a sense, it gives new meaning to GIGO -- garbage in, genius out! The 'in' was ordinal input of words -- which is 'garbage' compared to the resulting ratio scale measures that are shown in the Actual Column.
Verbal judgments are often more appropriate when judging qualitative factors, and all important decisions have qualitative factors that must be evaluated.
You might expect that using "fuzzy" words like Moderate, Strong, Very Strong, and Extreme would introduce error — and you'd be right. However, when we add redundant judgments — that is, instead of only saying A is 2 times B and B is 3 times C, which, if we assume these estimates are correct, can logically be enough on its own to conclude A is 6 times C. However, those judgments were not necessarily accurate. So now if we also compare A against C directly, and judge A to be 5 times C — something valuable happens. Not knowing which of the three judgments is correct, we use all three, and call this redundancy. You made redundant judgments in the exercise above. Those were processed in a way that reduces the 'noise' in the verbal judgments, producing accurate ratio-scale results.