Beyond The Pair: Ranking A Full Sector With Paired Comparisons, And How Much You Can Trust It - Part II

Relative, absolute, and resilience rankings for semiconductor firms often sharply diverge.

depositphotos_11100965-stock-photo-microchip.jpg
Source: DepositPhotos

<< Read More: Absolute Vs. Relative Valuation: What Happens When You Run Both On The Same SEC Data - Part I

1. The open question from Part 1

Part 1 of this series separated two questions that routinely get treated as one:

  • Relative valuation: is Company A cheap or expensive compared to Company B?

  • Absolute valuation (DCF): what is Company A worth on its own, from its own projected cash flows?

Running both on real duels (NVDA vs AMD, GOOGL vs META, AMZN vs WMT), plus a third lens — our Fundamental Resilience Index (FRI), which scores balance-sheet and cash-flow quality independent of price or growth — showed that the "winner" by one method is not reliably the winner by another. Three different, legitimate questions; three different, legitimate answers.

We also logged an honest limitation in that piece (section 2.4): a duel only ever compares two companies. Classical relative valuation compares a company to a peer group, not to one arbitrary rival picked for a headline. A pairwise "winner" can look completely different against a different opponent.

This piece is a direct attempt to close that gap — and, just as importantly, to measure how much it actually closes it, rather than simply asserting that it does.

2. What we expected to find, stated before we looked

Before running the numbers, we wrote down three testable expectations, so the results below can be checked against them rather than fitted to them after the fact:

  • H1 — Persistence of divergence. If the relative/absolute/resilience split from Part 1 was a real property of the data (and not an artifact of picking one specific pair), it should reappear at sector scale: several companies should rank near the top on one lens and near the bottom on another.

  • H2 — Resilience is orthogonal. FRI measures something structurally different (balance-sheet quality) from growth-and-margin factors (relative score) or cash-flow-implied value (DCF). We expected its ranking to correlate weakly, not strongly, with the other two.

  • H3 — Small samples are noisy, even with a "principled" model. A 2-round pilot gives each company only 2 data points. We expected the resulting rank order — however mathematically sophisticated the aggregation method — to be considerably less stable than it looks, and we designed a way to quantify exactly how unstable, rather than just gesture at the word "pilot."

All three held up, including in ways we did not fully anticipate (see section 6).

3. Method

3.1 Tournament design

Running a full round-robin across 12 companies (every company vs. every other) would require 66 duels — too many for a content pilot. We used the circle method, the standard algorithm for scheduling partial round-robins in sports and psychometrics, to build a 2-round schedule where the resulting "who played whom" graph is a single connected cycle touching all 12 companies (SWKS was dropped from the original 13-name group — smallest index weight, least distortion from removing it).

Round 1: NVDA–ON, AVGOMCHP, MUMPWR, AMD–MRVL, INTCQCOM, TXNADI Round 2: NVDA–MCHP, ON–MPWR, AVGO–MRVL, MU–QCOM, AMD–ADI, INTC–TXN

Connectivity matters here, not as a formality: Bradley–Terry strength estimates propagate through the graph of who-beat-whom, so if the graph split into two disconnected clusters, the model would have no basis to compare a company in one cluster to a company in the other. A single connected cycle is the minimum structure needed to rank all 12 names against each other, even though most pairs never played directly.

3.2 Aggregating win/loss into a ranking

From the 12 results we computed two relative-strength measures:

  • Copeland score — wins minus losses. Transparent, but coarse: with only 2 games each, most companies land in one of three buckets (+2, 0, or −2), producing wide ties.

  • Bradley–Terry (BT) rating — the standard model for paired-comparison data (1952; mathematical ancestor of the Elo system). Under BT, the probability that i beats j is s_i / (s_i + s_j), where s_i is i's estimated strength. We fit strengths by iterative maximum-likelihood (the MM/Zermelo algorithm), which lets strength propagate transitively: beating a team that itself beat a strong team counts for more than beating a team with a blank record, even with no direct comparison between the two indirectly-linked companies.

Two records in this dataset are "perfect" in the mathematical sense that matters for BT: a company that wins every game it played has, without correction, an estimated strength of infinity; a company that loses every game has an estimated strength of zero. Both break the model. We used a standard regularization — a small number of fictitious games against a league-average opponent, added to every company — which keeps every estimate finite. This is not a cosmetic fix; it is the mechanism that makes the ranking possible at all with this little data, and it means the exact numerical gap between adjacent ranks should not be over-read (more on this in section 6).

We then added two rankings that come from each company's own filings, independent of who it was paired against:

  • DCF-implied valuation multiple — Enterprise Value ÷ Revenue, an output of the DUEL DCF engine for each company individually (verified identical across both of a company's duel reports, confirming it does not depend on the opponent). Higher means the model's cash-flow engine assigns more enterprise value per dollar of revenue, given that company's own growth, margin, and ROIC inputs. This is not "upside vs. market price" — no live market price is used — it is a cross-sectional read of which cash-flow profiles the model rewards.

  • Fundamental Resilience Index (FRI) — our balance-sheet-quality composite, also company-specific.

4. Results

4.1 Round-robin outcomes

Duel

Score

Winner

NVDA vs ON

80–20

NVDA

AVGO vs MCHP

100–0

AVGO

MU vs MPWR

27–73

MPWR

AMD vs MRVL

73–27

AMD

INTC vs QCOM

12–88

QCOM

TXN vs ADI

57–43

TXN

NVDA vs MCHP

100–0

NVDA

ON vs MPWR

12–88

MPWR

AVGO vs MRVL

100–0

AVGO

MU vs QCOM

31–69

QCOM

AMD vs ADI

34–66

ADI

INTC vs TXN

8–92

TXN

4.2 Four rankings, side by side

Company

Record

Copeland

BT rank

DCF EV/Revenue

DCF rank

FRI score

FRI rank

TXN

2-0

+2

1

1.90x

9

60 (Moderate)

11

QCOM

2-0

+2

2

3.54x

5

89 (High)

3 (tie)

MPWR

2-0

+2

3

5.27x

3

98 (High)

2

NVDA

2-0

+2

4

19.84x

1

64 (Moderate)

10

AVGO

2-0

+2

5

16.06x

2

75 (High)

8

ADI

1-1

0

6

3.87x

4

87 (High)

6

AMD

1-1

0

7

3.12x

7

100 (High)

1

INTC

0-2

-2

8

n/a (negative EV)

12

32 (Low)

12

MU

0-2

-2

9

0.64x

11

76 (High)

7

ON

0-2

-2

10

1.93x

8

88 (High)

5

MCHP

0-2

-2

11

1.71x

10

73 (High)

9

MRVL

0-2

-2

12

3.45x

6

89 (High)

3 (tie)

5. Where the three lenses disagree

TXN won both its duels and holds the #1 BT rank, but sits at #9 of 12 on the DCF multiple and #11 on resilience (flagged for elevated leverage — D/E near 0.9, thin cash-to-debt coverage). The best relative-valuation record in the tournament belongs to one of the weaker absolute-value and balance-sheet stories in the group.

MRVL is close to a mirror image: last on BT, but tied for #3 on resilience (FRI 89) and mid-pack (#6) on the DCF multiple. A relative-only read tells you almost nothing accurate about MRVL's balance sheet.

AMD sits in the middle on both relative record (1-1) and DCF multiple (#7), but has the single best resilience score in the sector (FRI 100 — no flagged conflicts, full marks on every sub-metric). None of that is visible from who AMD beat.

NVDA and AVGO are the two names where the DCF layer agrees with the relative layer (#1 and #2 on both) — but resilience diverges. NVDA drops to #10 on FRI, driven by a Liquidity Runway metric well below the "strong buffer" threshold: a real trade-off between growth/valuation strength and balance-sheet cushion, not noise.

INTC is the control case: last or near-last on every lens, and the only company where the DCF engine returns a negative enterprise value outright (negative free cash flow breaks the model rather than merely lowering the number). When the underlying picture is genuinely weak across the board, all three methods agree — which is itself informative: it suggests the divergence seen elsewhere is signal about genuinely mixed fundamentals, not the model disagreeing with itself at random.

6. How much should you trust the #1 spot?

Before treating "TXN ranks #1" as a finding about TXN, it's worth being precise about what is actually driving that number.

TXN's two wins are against INTC (0-2, no wins anywhere in the network) and ADI (1-1 — its only credential is a single win over AMD). In a Bradley–Terry model, a win over a team that has itself beaten someone counts for more than a win over a team with a blank record, because strength propagates transitively through the graph. NVDA, AVGO, QCOM, and MPWR — the other four 2-0 teams — each beat two opponents that were 0-2 with no wins anywhere in the network. TXN is the only 2-0 team whose second win came against an opponent (ADI) carrying a live win of its own. That single indirect link is what separates TXN's #1 BT rank from a four-way tie at the top.

That is not a bug in the method — it is exactly what Bradley–Terry is designed to do. But it is also exactly the kind of result that should make you nervous about reading too much into a rank based on 2 games per company. To quantify that nervousness rather than just assert it, we ran a Monte Carlo check: simulate 12 companies with random "true" strengths, run the same circle-method schedule, and ask how often the Bradley–Terry #1 and the simple Copeland #1 disagree, purely from sampling variation, at different schedule lengths (4,000 simulated tournaments per data point):

Rounds per company

BT vs. Copeland disagree on #1

2 (this pilot)

64.4%

4

34.4%

6

25.2%

11 (full round-robin)

4.9%

At 2 rounds, which method you use for "who's #1" is close to a coin flip in terms of whether it agrees with the simpler count-based method — meaning the sophistication of Bradley–Terry is not yet buying you a stable answer at this sample size, only a differently unstable one. That instability is a property of the sample size, not of TXN specifically, and it is exactly why we're presenting this as a pilot finding to be revisited with more data, not a confirmed sector-leader claim.

7. How many rounds would it actually take?

Section 6 shows the #1 spot is shaky at 2 rounds. The natural next question — and the one this piece set out to answer quantitatively rather than qualitatively — is: how many rounds would be enough, and what do you gain at each step?

We simulated round-robin tournaments of the same 12-company circle-method design, at schedule lengths from 1 round up to the full 11-round round-robin (every company plays every other exactly once), with companies assigned random "true" strengths in each of 3,000 simulated tournaments per row. We then measured how closely the resulting Bradley–Terry ranking matched the (known, in simulation) true strength order:

Rounds per company

Total duels (12 cos.)

Rank correlation with true order (Spearman)

Top-3 identification rate

% of companies with a "perfect" record

1

6

0.32

37.0%

100%

2 (this pilot)

12

0.47

48.5%

56.7%

3

18

0.55

51.1%

34.5%

4

24

0.60

55.9%

22.8%

5

30

0.65

59.4%

16.1%

6

36

0.68

60.6%

11.2%

8

48

0.74

64.2%

6.4%

11 (full round-robin)

66

0.78

66.9%

3.1%

Two things stand out. First, the pilot design used here (2 rounds) sits at a rank correlation of roughly 0.47 with the "true" order — useful for spotting broad tiers, not reliable for claiming a specific #1. Second, more than half the field (56.7%) finishes a 2-round tournament with a mathematically "perfect" 2-0 or 0-2 record — which sounds decisive but, per section 6, is exactly the condition that makes individual rank order unstable, since a perfect record contains no information about how much better a company is, only that it won every game it happened to play.

Reading the table as a planning guide for future tournaments of this kind:

  • 1–2 rounds: cheap, fast, good for flagging broad tiers (clear winners vs. clear laggards) in content built around a pilot or a single publishing cycle. Individual placement within the top or bottom tier should not be treated as settled.

  • 3–4 rounds: meaningfully better (correlation ~0.55–0.60) at moderate extra cost (18–24 duels); a reasonable minimum if the ranking itself, not just the divergence-with-DCF story, is the point of the piece.

  • 5–6 rounds: a workable middle ground (correlation ~0.65–0.68) between full-groupwise rigor and the production cost of running dozens of duels.

  • 8+ rounds, up to the full 66-duel round-robin: diminishing returns per additional round, but the only regime where you can defend a specific numbered rank (as opposed to a tier) with real confidence.

8. Does this resolve the critique from Part 1?

Structurally, yes, in part: aggregating 12 pairwise duels into a connected network, and scoring every company against the implied strength of the whole group rather than one arbitrary opponent, is a genuine methodological upgrade over a single duel — it is what classical relative valuation's "compare to a peer group" actually requires, translated into paired-comparison statistics.

Empirically, the answer is more qualified, and the table in section 7 is the honest version of that qualification. The specific 2-round design used here reduces — but does not retire — the original criticism. A company's position in this pilot's ranking is more informative than a single duel's winner, but per sections 6 and 7, it is still closer to a coarse tiering than to a precise, stable rank. The critique from Part 1 (section 2.4) is best read now as: "resolved in kind, not yet in degree" — the method for comparing against a peer group exists and works; running it at a sample size that makes individual placements trustworthy is a separate, quantifiable cost, shown above.

9. Conclusions

  1. The relative/absolute/resilience divergence from Part 1 is not an artifact of picking one pair. It reappears at sector scale: TXN (#1 relative, #9 DCF, #11 resilience) and MRVL (#12 relative, #6 DCF, #3 resilience) show 6–8 rank-place gaps between lenses that measure genuinely different things.

  2. Resilience is the most independent of the three lenses, as expected (H2). AMD's #1 resilience score coexists with a middling relative record and DCF multiple; NVDA and AVGO's shared #1/#2 relative-and-DCF standing coexists with resilience ranks of #10 and #8. A high score on growth-and-margin factors or on a DCF cash-flow multiple tells you close to nothing about balance-sheet cushion.

  3. The one case with no divergence (INTC) is also the one case of unambiguous distress — last or near-last on every lens, and the only company where the DCF engine's output is degenerate (negative enterprise value) rather than merely low. This is evidence that the divergence seen elsewhere reflects genuinely mixed underlying pictures, not noise in the scoring methods themselves.

  4. A 2-round pilot is directionally useful and individually unstable, and this is now measured rather than asserted. Simulation shows a ~0.47 rank correlation with "true" order at this design size, a 64.4% chance that a simpler counting method (Copeland) would have crowned a different #1 purely from sampling variation, and 56.7% of the field finishing with a "perfect," information-poor 2-0 or 0-2 record. The #1 spot in this pilot (TXN) rests on a single indirect link in the comparison graph (its win over ADI, whose own credential is a single win over AMD) — a textbook illustration of exactly this instability, not a defect specific to TXN.

  5. The pairwise-vs-peer-group critique from Part 1 is structurally addressed but only partially retired in practice. The method for aggregating many duels into a group-relative ranking exists and works; per the table in section 7, making individual rank placements (not just tier placements) trustworthy requires roughly 5–8 rounds (30–48 duels) rather than 2 — a concrete, reusable number for scoping the next tournament of this kind, rather than a vague call for "more data."


Data source: SEC EDGAR 10-K/10-Q filings, retrieved via the DUEL platform (duelstocks.com). All rankings, scores and DCF outputs are generated by disclosed, fixed formulas from public filings — no analyst judgment or subjective inputs. Simulation results in sections 6–7 are based on synthetic tournaments with randomly assigned strengths, used only to characterize the statistical properties of the scheduling method, not to make claims about any specific company. This article is for informational and educational purposes only and does not constitute investment advice or a recommendation to buy or sell any security. Past data does not guarantee future results.

STOCKS IN THIS ARTICLE

Also Mentions:

Comments