
Methodology deep-dive, real numbers from real duels, computed and verified from scratch. Not investment advice.
Our first piece in this series flagged something worth revisiting: the weights behind our relative "duel" score — Revenue Growth 16%, Op. Cash Margin 15%, ROIC 15%, Sloan Ratio 12%, Asset Turnover 10%, Receivables Turnover 8%, Operating Margin 12%, FCF Margin 12% — are a reasonable expert starting point, not something backtested against realized returns. The tool lets any user override these with custom weights, which raises a question worth actually testing at three different levels of rigor: how much does the winner change if a reasonable person picks different, equally defensible weights?
Level 1: how the score works, and a clean fragility number
For each of the eight metrics, the two companies are compared and the higher-scoring one (lower, for Sloan Ratio, which is inverted) takes the entire weight of that metric — winner-take-all per metric, not graded by the size of the gap. The final score is the sum of weights attached to whichever metrics each company wins. We validated this by re-deriving the exact score for all 8 duels published to date from raw metric values — every one matched exactly.
That structure gives a clean, model-free fragility measure for free: moving weight Δ from any metric the winner won to any metric the loser won changes the score gap by 2Δ, so the minimum reallocation needed to flip a duel is simply gap ÷ 2.
Duel | Score | Gap | Weight-points to flip |
|---|---|---|---|
92:8 | 84 | 42.0 | |
10:90 | 80 | 40.0 | |
55:45 | 10 | 5.0 | |
32:68 | 36 | 18.0 | |
61:39 | 22 | 11.0 | |
73:27 | 46 | 23.0 | |
30:70 | 40 | 20.0 | |
75:25 | 50 | 25.0 |
AMZN vs WMT is the clear outlier — fragile by an order of magnitude compared to everything else in the dataset.
Level 2: Monte Carlo — does randomness agree with the algebra?
The gap÷2 number is exact but static — it doesn't say how likely a flip is if weights are drawn from a realistic range rather than deliberately engineered. We ran 10,000 simulations per duel, each drawing every weight independently from ±50% of its default value, then renormalizing to sum to 100% — a standard perturbation approach for testing sensitivity in multi-criteria scoring systems.
Duel | Gap | Monte Carlo flip rate (10,000 draws) |
|---|---|---|
NVDA vs AMD | 84 | 0.0% |
GOOGL vs META | 80 | 0.0% |
AMZN vs WMT | 10 | 14.8% |
CRM vs NOW | 36 | 0.0% |
VRTX vs REGN | 22 | 1.3% |
GFF vs WOR | 46 | 0.0% |
BHE vs PLAB | 40 | 0.0% |
KO vs PEP | 50 | 0.0% |
The simulation and the algebra agree completely: every duel with a gap ≥36 flipped in zero of 10,000 random draws. The two narrower duels (gap ≤22) are the only ones that ever flip, and the rate scales with how close the gap is to zero. That agreement is itself worth noting — a purely combinatorial shortcut (gap÷2) predicted the empirical simulation result without running a single random draw.
One more thing the simulation reveals that the algebra alone can't: which metric tends to be doing the flipping. Looking specifically at the 14.8% of AMZN-vs-WMT simulations where the winner changed, we can ask which factor happened to carry the most weight in those specific draws:
ROIC was the largest weight in 64.6% of flips
Sloan Ratio in 12.4%
Revenue Growth in 9.6%
Everything else, combined, under 14%
ROIC dominates the flip mechanism for this pair by a wide margin — a useful, specific answer to "what actually decides this one," not just "how likely is a flip."
Level 3: does a metric's weight actually equal its power? (Shapley-Shubik and Banzhaf indices)
Here's a subtler question the first two levels don't address at all: is a metric's stated weight the same thing as its real influence on who wins? In 1954, Lloyd Shapley and Martin Shubik asked exactly this question about weighted voting bodies (it's the same math later used to analyze voting power in the EU Council of Ministers and the UN Security Council) — and showed that nominal vote share and actual decisive power are frequently not the same number. A voting bloc can hold 20% of the vote and have almost no power to change any outcome, or hold 10% and be pivotal constantly, depending on how the other votes are distributed.
Our winner-take-all-per-metric structure is, mathematically, exactly this kind of weighted voting game: eight "voters" (the metrics), each with a fixed weight, deciding by simple majority (>50%) which side wins. So we computed the Shapley-Shubik index (and, as a cross-check, the related Banzhaf index) for every metric under the default weights — evaluating, across all 2⁸ = 256 possible ways the eight metrics could be split between two sides, how often each metric is the one that tips the outcome over the line.
Metric | Stated Weight | Shapley-Shubik Power | Power ÷ Weight |
|---|---|---|---|
Revenue Growth | 16.0% | 18.2% | 1.14x |
Op. Cash Margin | 15.0% | 16.1% | 1.07x |
ROIC | 15.0% | 16.1% | 1.07x |
Sloan Ratio | 12.0% | 11.8% | 0.98x |
Operating Margin | 12.0% | 11.8% | 0.98x |
FCF Margin | 12.0% | 11.8% | 0.98x |
Asset Turnover | 10.0% | 8.2% | 0.82x |
Receivables Turnover | 8.0% | 6.1% | 0.76x |
The pattern is real, if bounded rather than dramatic: bigger metrics carry systematically more decisive power than their stated weight implies, and smaller ones carry systematically less. Receivables Turnover is the clearest case — its published weight (8%) overstates its actual influence on outcomes more than any other metric's, and the same pattern (smallest weight, most overstated) held up when we re-ran the calculation under all four of DUEL's officially documented alternative weighting profiles (Growth, Quality, Value, Dividend & Stability). This isn't a flaw specific to this tool — it's a structural property of any winner-take-all weighted scoring system, first formalized 70 years ago in a completely different field (voting theory), and it applies here without anyone having designed it in.
Testing against the real, published investor profiles
DUEL documents four named weighting profiles for different investor styles. Applying all four to our 8 duels:
Duel | Default | Growth | Quality | Value | Dividend |
|---|---|---|---|---|---|
NVDA vs AMD | 92:8 | 95:5 | 95:5 | 85:15 | 92:8 |
GOOGL vs META | 10:90 | 8:92 | 8:92 | 18:82 | 10:90 |
AMZN vs WMT | 55:45 | 67:33 | 44:56 (flip) | 40:60 (flip) | 60:40 |
CRM vs NOW | 32:68 | 37:63 | 17:83 | 32:68 | 40:60 |
VRTX vs REGN | 61:39 | 64:36 | 56:44 | 65:35 | 56:44 |
GFF vs WOR | 73:27 | 72:28 | 74:26 | 80:20 | 66:34 |
BHE vs PLAB | 30:70 | 21:79 | 31:69 | 48:52 | 28:72 |
KO vs PEP | 75:25 | 72:28 | 84:16 | 72:28 | 71:29 |
Only AMZN vs WMT flips — the same duel every other method in this piece independently flagged as the fragile one. BHE vs PLAB is worth naming separately: it compresses from a 40-point blowout to a near-coin-flip (48:52) under the Value profile, without technically flipping, because BHE specifically wins the two metrics — Asset Turnover and Receivables Turnover — that the Value profile concentrates weight on.
What this actually resolves, and what it doesn't
The original limitation was that the base weights are a heuristic choice, not a backtested one. None of the three methods here answer whether 16/15/15/12/10/8/12/12 predicts future returns better than any alternative — that calibration question stays open. What we can now say with three independent, mutually-agreeing methods: the base weighting is not fragile in the way that would matter for practical use. A wide duel score (gap ≥30-ish) survived every random perturbation we threw at it and every real, named alternative profile. A narrow score is exactly where the three methods converge on the same warning — and, thanks to the Shapley-Shubik layer, we now also know that the underlying scoring structure systematically under-credits smaller metrics like Receivables Turnover regardless of which duel you're looking at. That's a genuinely new, structural finding about the mechanism itself, not just about any one comparison — the kind of result you can only get by testing the model's foundations with tools built for exactly this question, rather than testing weights one profile at a time.
References
Shapley, L. S., & Shubik, M. (1954). "A Method for Evaluating the Distribution of Power in a Committee System." American Political Science Review, 48(3), 787–792.
Banzhaf, J. F. (1965). "Weighted Voting Doesn't Work: A Mathematical Analysis." Rutgers Law Review, 19, 317–343.
Sloan, R. G. (1996). "Do Stock Prices Fully Reflect Information in Accruals and Cash Flows about Future Earnings?" The Accounting Review, 71(3), 289–315.
Saaty, T. L. (1980). The Analytic Hierarchy Process. McGraw-Hill.
Damodaran, A. (2012). Investment Valuation: Tools and Techniques for Determining the Value of Any Asset (3rd ed.). Wiley.
If you want to run any of this yourself — swap in your own weights, or test a pair we haven't covered — the live tool is free to try at duelstocks.com; every number in this piece traces back to a report you can pull there directly.
Not investment advice. All data sourced from public SEC EDGAR filings. Disclosure: I hold a long-standing position in NVDA (~5 years); no position in any other stock mentioned.



Comments
Log in or sign up to join the conversation.