
The question
DUEL's Resilience Report (STR) checks every company against three internal-consistency pairs: does revenue growth actually convert into free cash flow (Growth vs FCF), does a strong ROIC hold up once you check accrual quality (ROIC vs Sloan), and does the balance sheet's debt load actually match its cash cushion (Debt vs Cash). Any pair that fails triggers a score penalty and a flagged note.
Oracle (ORCL)'s ROIC-vs-Sloan flag — a headline ROIC of 84% sitting next to a Sloan Ratio signaling possible earnings inflation — was compelling enough to build a whole piece of coverage around. But one dramatic case raises an obvious question: is that the normal way a company fails this model's consistency check, or is it the rare exception dressed up as the norm because it happened to be the one we covered?
Research question: across a real cross-section of companies, what is the base rate of each of the three consistency-matrix conflicts, and which one actually drives most resilience penalties?
This is a different statistical animal from our earlier PCA piece on metric redundancy. That study asked how continuous metrics correlate with each other. This one asks how often a binary pass/fail threshold gets triggered — a base-rate estimation problem, which calls for a different toolkit entirely.
Literature
Sloan (1996) is the theoretical anchor for why ROIC-vs-Sloan should behave as an early-warning signal in the first place: firms with high accruals relative to cash flow tend to see earnings quality — and eventually returns — disappoint. But the base-rate question sits in a different, related literature: how common are earnings-quality red flags in a general, unscreened population of public companies? Beneish (1999), building the original M-Score, worked with a deliberately unbalanced sample — 50 known earnings manipulators matched against 1,708 non-manipulators — precisely because true manipulators are rare in the wild; his design had to oversample them to have enough positive cases to model at all. Dechow, Ge, Larson, and Sloan (2011), studying SEC enforcement actions for accounting misstatement across 23 years, reach the same conclusion from actual population data: material misstatement is a low-base-rate event, not a coin flip. Our prior on the ROIC-vs-Sloan conflict, going in, should be "rare" — and the question is whether the DUEL data agrees, and whether the other two conflict types share that rarity or not.
For the statistics of estimating a rate itself: Cochran (1977) gives the standard sample-size formula for estimating a proportion with a target margin of error. But once the rate is estimated, especially near the boundaries (very low or very high), the ordinary normal-approximation ("Wald") confidence interval performs badly — it can even produce nonsensical negative lower bounds. Wilson (1927) proposed a score-based interval that stays well-behaved at extreme proportions, and Newcombe (1998) confirms it as the more reliable choice for small-sample or rare-event proportions like the ones this article deals with directly. All confidence intervals below use the Wilson method for that reason.
Data and method
Panel: the same 60 companies (30 duels) used in our prior PCA piece, now with each company's STR Resilience Report pulled instead of its Battle Report — same underlying population, different report artifact, which lets the two studies sit side by side on one consistent panel.
Variables: for each company, the three consistency-matrix outcomes as reported (OK / CAUTION / RISK), plus the overall FRI score, resilience tier (HIGH/MODERATE/LOW), base score, and penalty.
Method: base rate = count of non-OK outcomes ÷ 60, computed separately for (a) any non-OK flag (CAUTION or RISK combined) and (b) RISK specifically, for each of the three checks. 95% confidence intervals via the Wilson score method. Cochran's (1977) minimum-expected-cell guideline (≥5 events for a stable rate estimate) is applied explicitly to flag which of the three rates are trustworthy at this sample size and which aren't yet.
Results
Base rate per conflict type
Check | Non-OK (CAUTION+RISK) | Rate | 95% Wilson CI | RISK only | Rate |
|---|---|---|---|---|---|
Growth vs FCF | 2 / 60 | 3.3% | [0.9%, 11.4%] | 0 / 60 | 0.0% |
ROIC vs Sloan | 2 / 60 | 3.3% | [0.9%, 11.4%] | 1 / 60 | 1.7% |
Debt vs Cash | 26 / 60 | 43.3% | [31.6%, 55.9%] | 19 / 60 | 31.7% |
This is not a subtle difference. Debt vs Cash is triggered roughly thirteen times as often as either of the other two checks. Growth vs FCF and ROIC vs Sloan — the conflict type that inspired this whole line of inquiry — land at the exact same rate, 3.3%, tied for the rarest outcome in the entire matrix.
How many companies carry more than one flag
Simultaneous conflicts | Companies | Share |
|---|---|---|
0 (clean) | 33 | 55.0% |
1 | 24 | 40.0% |
2 | 3 | 5.0% |
3 (all checks fail) | 0 | 0.0% |
55% of companies in this panel pass all three checks cleanly. No company in the sample failed all three simultaneously — the theoretical worst case never showed up. Of the 27 companies (45.0%, 95% CI [33.1%, 57.5%]) carrying at least one flag, the overwhelming majority (24 of 27) carry exactly one, and it is almost always the same one.
What "Debt vs Cash" actually flags
Every RISK-level Debt vs Cash note in this panel follows the same template: a debt-to-equity ratio well above 1x combined with cash covering only a fraction of total debt (e.g., HD: D/E 3.3, cash covers 0.0x of debt; CAT: D/E 1.9, cash covers 0.1x). This is a leverage signal — how much debt sits on the balance sheet relative to the cash cushion available to service it — not an accounting-quality signal. It is mechanically far more common for a mature, capital-intensive business to run meaningful leverage than for its earnings to show acute accrual manipulation, and the 13x frequency gap in this data is consistent with that basic distinction.
Not every LOW-resilience company has a conflict
Seven companies in this panel land in the LOW resilience tier. Only three of them (STX, HD-adjacent cases, SHW/KO with Debt vs Cash RISK, STX with two simultaneous flags) got there via a consistency-matrix penalty. The other four — INTC, OXY, SNDK, and one more — arrived at LOW resilience with a weak base score and zero conflicts: every individual component (Cash-to-Debt Coverage, Investment Self-Sufficiency, Earnings Quality, Liquidity Runway) simply scored low on its own merits, with nothing contradicting anything else. That's a meaningfully different situation from Oracle's case: "consistently weak across the board" and "internally contradictory" are two different failure modes, and this panel shows both occur, with the contradiction-driven failure being the less common of the two.
Discussion
The finding inverts the intuition that motivated this whole project. The Oracle-style flag — a dramatic-looking metric next to an earnings-quality red flag — is genuinely rare, at the same 3.3% rate as the other "quality" check (Growth vs FCF). It's also, not coincidentally, the most narratively interesting kind of flag, which is exactly why it made for compelling video coverage in the first place. But the conflict that's actually driving nearly all resilience penalties in a broad cross-section of large, well-known public companies is the least dramatic one: ordinary balance-sheet leverage. Beneish's (1999) oversampling of manipulators to get a workable dataset, and Dechow et al.'s (2011) population-level confirmation that misstatement is rare, both point the same direction our data does: the dramatic accounting-quality story is the tail event, not the norm — and a model built to catch it should expect to flag it rarely, which is exactly what happened here.
Practical takeaway for DuelStocks
This changes how the Resilience Report's own flags should probably be talked about, and it's a genuine content opportunity, not just a research footnote:
Reframe the marketing angle. The Oracle-style "hidden earnings inflation" story is a great hook, but it's the exception, not what the model is mostly catching. The honest, and arguably more useful, positioning is: "Our model mostly catches ordinary leverage risk — the kind that shows up in roughly 4 out of 10 companies we check — not just rare accounting-quality drama." That's a stronger trust signal than implying every flag is an Oracle-caliber story, and it's backed by this data rather than by one case.
A calibration line in the STR report itself — something like "Debt vs Cash flags occur in ~43% of companies we've analyzed; ROIC vs Sloan flags occur in ~3%" — would let a reader immediately judge whether a given flag is common or genuinely unusual, instead of reading every CAUTION/RISK note as equally alarming. This is a product suggestion worth considering, not a change being made based on one pilot.
A natural next Short: "The Real Reason Most Resilience Penalties Happen (Hint: It's Not Accounting Fraud)" — the 43% Debt vs Cash statistic is a strong, honest, non-clickbait hook that's directly defensible with real numbers, unlike a generic "watch out for red flags" video.
Limitations
Two of the three rates are not yet reliably estimated. Growth vs FCF and ROIC vs Sloan each have only 2 observed events in 60 companies — below Cochran's (1977) rule-of-thumb minimum of 5 for a stable rate estimate. Their Wilson intervals (both [0.9%, 11.4%]) are wide relative to their point estimate, and either metric's true population rate could plausibly be twice or half of what we observed here. The Debt vs Cash estimate, with 26 events, is on much firmer ground.
Cross-sectional, one snapshot. Whether these base rates hold up a year from now, or across a different set of 60 companies, is untested.
Sector composition matters and isn't controlled for. This panel leans toward large, established public companies; a panel weighted more toward smaller or newly public companies might show a different leverage/accrual mix entirely.
Same panel as the PCA piece, on purpose — this lets the two studies be read together, but it also means neither study is a fully independent replication of the other's population.
None of this changes DUEL's actual scoring or penalty logic in the product. It's a base-rate finding about how often each flag fires in the real world, not a proposal to alter the underlying thresholds.
Bottom line
Of three consistency checks in DUEL's resilience model, one — Debt vs Cash — fires in roughly 4 of every 10 companies checked. The other two, including the exact type of flag that inspired this research (ROIC vs Sloan), each fire in about 1 in 30. The dramatic story is the rare one. The common one is ordinary leverage — which is, in its own way, the more useful thing for a resilience check to be catching reliably.
As always: every duel's full Resilience Report, with every raw number behind these flags, is available at duelstocks.
References
Beneish, M. D. (1999). The detection of earnings manipulation. Financial Analysts Journal, 55(5), 24–36.
Cochran, W. G. (1977). Sampling Techniques (3rd ed.). Wiley.
Dechow, P. M., Ge, W., Larson, C. R., & Sloan, R. G. (2011). Predicting material accounting misstatements. Contemporary Accounting Research, 28(1), 17–82.
Newcombe, R. G. (1998). Two-sided confidence intervals for the single proportion: comparison of seven methods. Statistics in Medicine, 17(8), 857–872.
Sloan, R. G. (1996). Do stock prices fully reflect information in accruals and cash flows about future earnings? The Accounting Review, 71(3), 289–315.
Wilson, E. B. (1927). Probable inference, the law of succession, and statistical inference. Journal of the American Statistical Association, 22(158), 209–212.
A DuelStocks methodology deep-dive. Not investment advice. All data sourced from public SEC EDGAR filings.



Comments
Log in or sign up to join the conversation.