Model Register
One entry per model: who owns it, what it assumes, what validates it, what monitors
it, and when it was last reviewed. Shaped after OSFI Guideline E-23 and kept honest by a test —
tests/parity/test_model_register.py compares every constant here against
the engine's source and fails the build when they disagree, so this page cannot quietly drift from
the code it describes.
It records what is not true, because that is the only reason a register is
worth having. Independent validation, in the E-23 sense of a validator separate from the developer,
is not available here: this is a single-developer product, and where a model is
checked it is checked by an independent oracle — a test written from the statute or the
stated law rather than from the implementation — never by an independent person.
Provenance is graded, not asserted. An earlier version of this page carried one free-text
“source” field, and ten entries had prose beside them. A reader could not tell a citation from an
excuse — one entry's “source” read “no citation”, and another pointed at a second constant
that was itself ungrounded. Those ten are now graded, and six of them do not survive.
external source a named third party a reviewer can look up ·
derived / governed traceable to governed config ·
recorded decision a human choice with a rationale and a date — justified under E-23
NOT a source text that explains what the constant does, not why its value is that value ·
nothing written down
Only the first three count as justified. The figure below is computed
from the grades, not from how many entries have words next to them.
67 / 75
assumptions with any justification at all
9 / 14
not monitored against outcomes
2
entry points with no law-stating gate
register version 2.0.0 · generated 2026-09-23
· grades: external 6 · derived 1 · decision 60 · mechanism 6 · none 2
Fair value (DCF + ensemble)
fair_value · owner vishalsinha · version 2.8.0 · last reviewed 2026-09-24 · valuation/models/dcf.py
Monitoring
NOT MONITORED on a schedule valuation/validation.confidence_calibration computes the realised median error by confidence tier on demand from stored scans (closing prices as stored, so a split between scans distorts it) and the Strategy advisor quotes it; no graded record is published
Metric: fair-value error vs realised price. Grader: none. Log: none.
Validation
invariant tests/invariants/test_dcf_discount_invariants.py (no price feedback into the discount rate), test_model_entry_point_invariants.py (WACC floor survives any leverage; a debt-free firm gets exactly its floored cost of equity), and test_ensemble_confidence_invariants.py for valuation/aggregate.blend, which is this model's aggregation step but lives outside its file so it is not listed above: the top tier now requires three corroborating families, and a fully price-rejected name suppresses rather than guesses
Assumptions
Limitations
- 2.8.0 (2026-09-24): THE CONFIDENCE TIER IS MEASURED BEFORE THE PRICE ANCHOR, NOT AFTER. blend discards any model outside [0.33x, 3.0x] price and then measured dispersion over the SURVIVORS, so the tier reported the agreement it had manufactured by removing the models that disagreed. Measured on the 2026-09-23 scan of 2,891 rows: 1,031 of the 2,035 rows carrying a published tier carry a different one on the pre-anchor basis, and 85 of the 216 HIGH labels were not HIGH. THE THRESHOLDS MOVED WITH THE BASIS AND HAD TO: CV_HIGH/MEDIUM/LOW were calibrated when the median post-anchor dispersion was 0.267 (p90 0.466, n=1,997); the pre-anchor median is 0.456 (p90 0.822, n=2,620), which put the old 0.50 suppression cut-off near the middle of the distribution instead of in its tail and would have withdrawn 840 of the 2,035 published values - 41% of the book - as an artifact of an unmoved constant. Re-fitted BY SWEEPING the live engine rather than by scaling the old numbers (a first pass at matching quantiles over-suppressed, because most of today NOISE comes from causes other than dispersion): CV_HIGH 0.15->0.20, CV_MEDIUM 0.30->0.50, CV_LOW 0.50->0.80. Coverage lands at 1,997 rows against 2,035, with 137 appearing and 175 withdrawing. THE CHANGE IS THE RANKING, NOT THE COUNT: 1,168 of 2,891 tiers move, HIGH is 200 against 216, NO POINT ESTIMATE MOVES AT ALL (0 of 2,891 fair values change), and the labels now sit on the rows whose models genuinely agreed - AAPL and RY.TO fall HIGH to MEDIUM, MSFT MEDIUM to LOW, NVDA LOW to MEDIUM. A CLAMPED JUSTIFIED MULTIPLE NO LONGER CORROBORATES: a rail is a constant and agrees with whatever it is placed beside, so it still informs the point estimate but bars HIGH, which costs 18 of the 218 rows that would otherwise reach it (78 of 2,891 rows carry at least one clamped family vote). Both dispersions and the clamped-vote count now ship on EnsembleData so a reader can audit the label. These cut-offs are a HOUSE CALIBRATION fitted to this book; no external convention sets them, and the register records that rather than implying one.
- 2.7.0 (2026-09-24): WHAT EACH MODEL CLAIMS, AND WHOSE MONEY IT IS. All 21 models now declare the convention their NAME invokes - exact, proxy with a note saying what they really compute, or house where none is claimed (6/8/7). Three were misnamed: EPV capitalises normalised EPS at the cost of EQUITY where Greenwald capitalises normalised after-tax EBIT at the WACC and then adjusts for debt and cash; Graham Revised described its denominator as the AAA corporate yield and used the government 10-year, reading about 18% high for a measured median blend effect of -0.67%; Residual Income damps its excess-return term by a house persistence of 0.5 that is no part of the standard formula. A proxy is now LABELLED ON THE PER-MODEL PANEL, where the model name appears, not only in this register. The model formerly called EV/Rate-base was RENAMED to EV/Net PPE: a rate base is a commission determination this engine does not have, and the model values a utility off net property, plant and equipment - the name now states the quantity computed and claims nothing. COMMON VS TOTAL: wherever one side of a ratio belongs to the common shareholder the other now does too. ROTE was net income BEFORE preferred over tangible COMMON equity (of 201 SPREAD banks publishing a fair value, 59 carry preferred; overstated a median 3.6%, 5%+ on 14, and 220% on FLG); the payout behind sustainable growth divided the COMMON dividend by TOTAL earnings per share (64 of 287, retention overstated a median 0.8%, 5%+ on 4); and REIT FFO per share divided the Nareit TOTAL by the COMMON share count (22 of 119 REITs, 21 publishing, overstated a median 1.8% and up to 20.1%). One owner now performs the subtraction. PUBLISHED EFFECT IS SMALL AND THAT IS THE HONEST NUMBER: 4 of 2,891 fair values moved on the ROTE and payout fixes (AGM -5.0%, VLY -2.4%, WFC -2.2%) and 25 on the FFO fix, because the justified multiple is clamped to [0.5, 3.0] and 24 of the 61 affected banks sit ON a rail where a ROTE change cannot move the published figure, with the blend damping the rest. STILL NOT FIXED, BY DECISION: the EV-to-equity bridge subtracts debt and adds cash only, leaving preferred and non-controlling interests with the common shareholder on 167 of the 2,106 rows that publish. No preferred or NCI BALANCE exists in the schema - only the dividend flow - and capitalising a flow at an assumed rate would invent a number where a visible omission is more honest. It is published in inputs_used on every affected row.
- 2.6.0 (2026-09-24): PROJECTED GROWTH IS MEASURED, AND MAY POINT DOWN. The four upward biases the 2026-09-24 audit of this engine recorded are closed. The CAGR span is the years that ELAPSED rather than the count of profitable ones (a loss year between two good ones compressed the compounding window: 100, -50, -50, 121 read 21%/yr where three years passed and the rate is 6.6%; 1,155 of the 2,753 rows with two or more years of history had such a gap). Growth is no longer floored at zero, so a business whose cash flow falls every year is projected to SHRINK rather than valued flat forever (972 of 2,753 were shrinking). Damodaran, "Valuing Distressed and Declining Companies" (2009), pinned in docs/sources/, is explicit that perpetual growth "in some cases ... can even be negative"; the -10% BOUND, however, is a house judgement and is labelled one at MAX_DECLINE, because that paper prescribes no number. The base averages the last three years AS REPORTED, loss years included, rather than only the profitable ones. The 40/60 historical-to-analyst growth blend no longer ratchets against history, so a more cautious analyst view can lower it. And where no rate exists at all - a company that ended the period in a loss - the model DECLINES to value the row instead of substituting zero. Measured on the 2,891-name scan of 2026-09-23: 542 fair values changed, 133 were withdrawn, 32 appeared, 331 tiers moved, and 102 of 2,577 buy/sell verdicts changed. KNOWN LIMIT OF THAT MEASUREMENT: the offline scorer runs without the vendor feed, so the growth-blend change is not exercised by it and is proven by unit test instead.
- 2.5.0 (2026-09-24): THE DENOMINATION, THE COHORT AND THE CLASSIFIER, on the LIVE path. The FX rate between a row's reporting and listing currency is now STORED (fundamentals.reporting_fx_rate, schema v8) instead of re-derived as market_cap / market_cap_reporting: price_refresh rescales market_cap alone every two minutes in-session, so the implied rate tracked the SHARE PRICE and every reporting-currency input scaled with the quote. Measured with the real refresh arithmetic over the 165 snapshot rows that carry a rate, a +5% price move corrupted the rate on 165 of 165 before and 0 of 165 after. The stale-macro-cache fallback is now per domicile (US 4.3%, CA 3.83% - Government of Canada 10-year benchmark, Bank of Canada Valet BD.CDN.10YR.DQ.YLD, observation 2026-09-22); one static for both had silently erased the whole Canadian basis whenever the cache went stale. THE LIVE DEEP-DIVE PANEL was found assembling three of its own inputs by hand rather than reading the deciders: the denomination (112 of 2,891 rows valued at par against their own headline), the peer cohort (a fee-earning financial peered against all of Financial Services, median price to tangible book 1.734 over n=430 against 5.073 over its own n=82), and the classifier's fields - it carried five of the eight business_class reads and omitted every signal that identifies a FUND, so 36 of 2,891 rows the nightly scan suppresses entirely were being valued on the stock page. Two Canadian closed-end funds publishing at MEDIUM confidence were added to the reviewed classification list; the automatic investment-company test is American (SEC filings plus the Investment Company Act 3x line) and cannot see a Canadian fund.
- 2.4.0 (2026-09-23): ONE DEFINITION PER PUBLISHED QUANTITY. Margin of safety is (FV-price)/FV everywhere, restated from the divide-by-price form at every call site (0 of 6,240 verdicts moved - the restatement is algebraic). A withdrawn or untrusted estimate no longer alerts, narrates or carries a range. Which models a row may run has ONE owner (class_policy); the retired sector rule had been denying a fee-earning financial the whole cash-flow family. THE COST-OF-CAPITAL FLOOR IS 7.0%, NOT THE 8.5% THE 2.1.0 ENTRY BELOW RECORDS - it moved on 2026-09-18 and this register did not say so until now. The same 4.5-point spread that defines it (7.0% = 2.5% terminal growth + 4.5) now guards every model dividing by (cost of equity - growth). That cap is NOT one-directional: (ROE-g)/(k-g) is DECREASING in growth whenever a company earns less than its cost of equity, so capping growth RAISED the fair value of 103 of the 2,089 names valued both ways (median +11.3%, largest +247%) while lowering 553 of 2,089 (median -12.5%), and moved 107 of 2,577 buy/sell verdicts.
- 2.1.0 (2026-09-16): DISCOUNT-RATE BASIS (valuation/discount.py), switched by VALUATION_DISCOUNT_RATE=2026, default OFF. On: ERP = Damodaran's January-2026 implied mature-market premium 4.23% + country premium (US 0.23%, CA 0.00%); risk-free = the domicile's cached 10-year yield (US ^TNX, Canada BoC Valet), read from the macro cache and never fetched, legacy 4.3% when absent/stale/implausible; the equity models discount on 0.67 x industry levered beta + 0.33 x firm beta. The 8.5% cost-of-equity FLOOR stays and is the BINDING assumption for the low-beta classes (regional banks, utilities, P&C insurers, REITs compute to 5.3-7.1% and are floored) — the register records the floor, not the beta, as what sets their rate. The FCFF path keeps its sector unlevered beta and Hamada relever. This reverses the 2026-09-10 decision to keep the historical premium as a stable yardstick; the owner made that call on 2026-09-15 with the pro-cyclicality cost stated.
- 2.0.0 (2026-09-16): per-class VALUATION POLICY (valuation/class_policy.py) — each enabled business class runs the models that mean something for it at policy weights, on a class growth axis (banks/insurers: retention x ROE; REITs: FFO/share; utilities: rate-base growth; commodity: FCF capped at 3%), with seven class anchors (Justified P/TBV, Justified P/B, P/FFO Peer, FFO Perpetuity, EV/Net PPE, Normalized EV/EBITDA, EV/Revenue growth-adjusted). The default class is valued exactly as before (sector policy, unweighted blend). The dynamic P/E, P/B and P/S peer medians fire ONLY while a class is enabled (they had fallen to the static tables for every name since inception: the ratio divided market cap by a per-share figure; switched on alone they moved ordinary names' fair values by up to +200% on the prod copy, so they come on with the programme, reviewed, never on a deploy). Utility rate-base growth into the DDM is capped at 4% (6% against the 8.5% floor priced the class at 1.13x the market). ONE producer since 2026-10-07: the scan and every price refresh blend through ensemble_store (valuation/run.py), each row at its own close, and the live deep-dive serves that stored record instead of valuing the stock again on Yahoo inputs. Known: EV/rate-base uses net PPE, not the regulatory rate base; the class multiples fall to static values below 3 class peers.
- Single-currency: one ERP and one terminal growth rate for every domicile.
- PRICE-ANCHORED, and materially so (audit C3, measured 2026-09-10 over 1,009 production names): the sanity filter that drops models falling outside 0.33x-3x the current price bound for 822 of 1,009 names (81.5%), discarded 2,914 of 10,995 applicable model outputs (26.5%), and where it bound moved the blend a median 8.0% against the same models with the anchor off (p90 54.4%). It exists to reject broken inputs and it does that, but a quarter of all model output being discarded for disagreeing with price means the fair value is substantially a function of price. 114 names were suppressed entirely.
- Confidence rests on very few numbers: every published tier is computed over 2 to 4 family representatives, median 3. HIGH now requires 3 (fixed 2026-09-10) — before that, 139 of 221 HIGH estimates, 62.9%, rested on just two.
- The ERP is STATIC by design (see the assumption note), so fair values do not respond to the level of rates or to how expensive the market is. In an expensive market this tool will find fewer bargains than one using an implied premium — that is the intent, and the cost is that fair value lags a genuine repricing of risk.
- The discount rate was adjusted by a score containing valuation until C1 re-keyed it to quality pillars only.
- The point estimate is published beside a bull/base/bear scenario range (built from the company's own normalised cash flow) and an interactive DCF, but the blended figure itself carries no error band.
Entry points & executable coverage
Deep Score (sector + size relative)
deep_score · owner vishalsinha · version 2.1.3 · last reviewed 2026-10-06 · deep_score_engine.py
Monitoring
WIRED, IMMATURE grades matured cohorts; none have matured
Metric: sector+size-relative excess forward return by score decile. Grader: scripts/deep_score_track_record.py. Log: winner_predictions.
Population test
partly certified 2026-10-06 scripts/backtest_deep_score_population.py · primary: 23,694 graded name-years, rank-IC +0.104, tercile spread +0.067, t 3.10, deflated Sharpe 0.992 (N=2), halves +0.073 / +0.110 -> certified; secondary: 22,184 graded name-years, rank-IC +0.106, tercile spread +0.073, t 2.89, deflated Sharpe 0.992, halves +0.065 / +0.130 -> not certified (scripts/backtest_deep_score_result.json; pre-registered df411c34, tag deep-score-preregistration-2026-10-06)
Scope: CERTIFIED: the 85-point fundamentals core (four business pillars) as a rank of the three-year sector+size-relative excess return over US 10-K filers FY2009-2022. NOT certified: the five-pillar variant (valuation on a public-float proxy, t 2.89). NOT tested: the Valuation pillar as shipped, PEG, the label cut-offs, TSX listings, the sizing bands, any probability.
Validation
population test (fundamentals core) + distribution invariant The 85-point fundamentals core's rank: certified by the pre-registered point-in-time population test, scripts/backtest_deep_score_population.py (result scripts/backtest_deep_score_result.json; 2.1.3) - the Valuation pillar, the label cut-offs and the bands are not covered by it. tests/invariants/test_deep_score_calibration_invariants.py — labels must land on the fractions their cliffs name; proven red against the pre-C5 behaviour. tests/invariants/test_deep_score_class_invariants.py — every class is scored on its own set in its own cohort, computes >= 70% of its factors on a six-class universe, the label mix still lands on the cliffs with every class on, classes-off is byte-identical, a thin class reverts to the default set; each gate proven red against its named injection (2026-09-16). tests/invariants/test_deep_score_bank_invariants.py — each balance-sheet class sits at its own pillar midpoint, one basis per cohort, a thin class reverts to the standard basis, every uncomputable substitute is None; each gate proven red against its named injection (2026-09-15)
Assumptions
Limitations
- 2.1.3 (2026-10-06): TESTED - THE FUNDAMENTALS CORE IS CERTIFIED AS A RANK, NARROWLY; THE MODEL IS UNCHANGED. CFA Standard V(A) asks that a model's output be tested before it is distributed. The test (scripts/backtest_deep_score_population.py) was pre-registered in its own docstring and committed (df411c34, kept reachable by the tag deep-score-preregistration-2026-10-06; rebased onto the integration branch with identical content) before any run read a return. POPULATION: every US 10-K filer FY2009-2022 in the SEC Financial Statement Data Sets, scored by the shipped engine as of each fiscal year from what it had filed at the time, graded on the three-year return from the first bar after its 10-K (a bankruptcy inside the window -100%) minus the mean of its own peer group and size tier that year; the size tier from the public float the filer states on its 10-K cover, the one market value in the record a split cannot distort. FUNNEL (primary): 65,318 as-of rows -> 40,916 scored, one vote per issuer (24,358 unscored: no history, a fund, or a pillar with nothing measurable) -> 23,726 graded (22,703 priced, 1,023 bankrupt; not graded: 11,248 with neither a price path nor an exit, 2,705 acquired and 2,474 gone dark with no price, 641 entries under $1, 65 other delistings, 57 open windows) -> 23,694 with a peer-group excess, 2,344 / 3,804 / 4,850 / 7,085 / 5,611 across the five three-year blocks. PRIMARY (the four business pillars, 85 of 100 points, under a research-only override of the pillar set - nothing in the engine changed): rank-IC +0.104; top-minus-bottom-tercile excess +6.7 points over three years; L/S Sharpe 1.39 over five independent blocks, t 3.10 against a bar of 3.0; deflated Sharpe 0.992 over the 2 trials run; held-out halves +0.073 (FY2009-2014) and +0.110 (FY2018-2022) -> CERTIFIED by the house certify() rule (base_rate_known passed as the Inflection program passes it: the Deep Score publishes no probability). A NARROW PASS ON ONE GATE and a MODEST EFFECT: t is 3.10 against 3.0, and two of fourteen years ran the wrong way (FY2011 -2.1 points, FY2022 -14.0). TIES: the published score is an integer, so names tie at the tercile edges, and the pre-registration did not specify how ties are broken; the recorded verdict uses the history's storage order (a stable sort over the rows as the database returns them). With t 3.099 that close to the 3.0 bar, the certification is a NARROW PASS and is published as one. Mean three-year excess by the label the four-pillar rank would carry (descriptive; not the shipped five-pillar label): STRONG BUY +3.1 points (n 4,081), BUY +2.1, HOLD +1.2, MARGINAL -0.2, AVOID -2.8 (n 8,624). SECONDARY (all five pillars, P/E and FCF yield from the float, PEG unknown): rank-IC +0.106, t 2.89 -> NOT certified. NOT TESTED, and no surface may say otherwise: the Valuation pillar as shipped (vendor P/E, PEG, FCF yield), the shipped composite's label cut-offs and fractions, TSX listings (not in the SEC record), the live vendor-sector cohorts (a SIC map stands in, and leaves 1,725 of the harvest's 13,986 filers - 12.3%, 8,494 of the 65,318 FY2009-2022 as-of rows - to its default Industrials / Conglomerates peer group, among them commercial biological research, laboratory instruments, coal and education; remapping them would be a new trial for all three engines, so the count is disclosed and pinned by tests/test_winner_recon.py instead), the per-share growth factors (withheld: 12-28% of those classes' rows carry a split-sized share-count jump the record cannot explain), the inputs the record does not carry (tangible equity, net-interest share, loss ratio, EBITDA, cash, dividends, analyst growth), the sizing bands, and any probability. KNOWN GAP: 917 FY2009-2022 rows of names that left the market have no delisted-name price for their filing date (the production-host job was not re-run) and are not graded. THE HARNESS WAS REPAIRED FIRST, and Winner Odds and the Inflection program were re-run on it (winner 1.16.1; inflection's population record): the SIC map read drug makers as chemicals and aircraft makers as retailers; the 2023q1 data set had never loaded, so 3,222 FY2022 rows carried a year-late filing date; and the as-of row dropped return on equity, growth, debt/equity and the class series while inventing a 0% payout for every REIT and utility. Every surface that shows a label or a band now says what was tested and what was not (tests/parity/test_recommendation_disclosure.py).
- 2.1.2 (2026-09-30): THE CLAIM IS CORRECTED; THE MODEL IS UNCHANGED. THE DEEP SCORE IS NOT YET TESTED AGAINST FORWARD RETURNS. Two statements were wrong. (a) deep_score_model.py and the Methodology said a historical test was IMPOSSIBLE (>= 5 independent windows from a ~10-year archive gives 3). Since 2026-09-27 the point-in-time population harness (scripts/backtest_winner_population.py, every US 10-K filer FY2009-2022, the harness Winner Odds and the Inflection Score were tested on) makes a population test of its fundamentals possible. That test is PLANNED and has NOT been run; no result exists, pass or fail. (b) 2.1.1 below said the 2026-10-12 grading of one cohort IS the V(A) test: one 126-day window in one regime, with no count of independent periods, is a diagnostic, not a test the house's own certify() rule would accept (CERT_MIN_PERIODS=5). Until the population test runs and is recorded here, every surface that renders a STRONG BUY / BUY / AVOID label or a sizing band says the labels have not been tested against later returns: the web InvestmentDisclaimer (both modes), SizingPanel, the global search, the landing score explorer, and the Android stock screen (label and band), enforced by tests/parity/test_recommendation_disclosure.py. The public /method page no longer calls the verdicts back-tested.
- 2.1.1 (2026-09-25): THE PRICE HISTORY NEEDED TO TEST THE GROWTH PILLAR NOW EXISTS; THE WINDOW HAS NOT CLOSED. The bulk OHLC backfill had never been run and, at period="max" for every ticker, would have written a projected 1.16 GB onto a production volume with 2.9 GB free. Bounded to 3y for universe names (holdings keep the full history the crisis replay needs) and run: price_history went from 105 of the scored names to 2,786 of 2,792, and 988 of the 1,059 names scored on 2026-06-08 now carry closes spanning that scan to today. The Growth pillar is therefore still disclosed rather than re-weighted, but the reason has changed from missing data to an unclosed window — that cohort grades at 126 days on 2026-10-12. CFA Standard V(A) requires a quantitative model's output to be TESTED before distribution; it does not require a live forward record, so the retrospective grading of this cohort IS the test.
- 2.1.0 (2026-09-24): THE GROWTH PILLAR DECIDES THE TOP OF THE RANKING, AND IT IS BUILT ENTIRELY ON HISTORICAL GROWTH. Measured on the production scan of 2026-09-24 over 2,577 scored names: the Spearman between the shipped score and the same score with the pillar removed is 0.9700, which understates it badly — 45 of the top 50 names (90.0%) are in that set ONLY because of the pillar, 6 of the 8 names the live screen surfaced were Growth-lifted (DLO by 207 places), and those 8 average 20.6 of 25 Growth points against a universe mean of 12.5. Both of the pillar's factors (revenue_growth w=12, earnings_growth w=13) are PAST growth, and the literature is against extrapolating it: Chan, Karceski & Lakonishok (Journal of Finance, 2003) find no persistence in long-term earnings growth beyond chance, and Lakonishok, Shleifer & Vishny (Journal of Finance, 1994) find the high-growth glamour end underperforms over five years. NOT RE-WEIGHTED, and deliberately: whether the pillar SIGNS wrong on this book is not testable YET — until 2026-09-25 price_history covered 127 tickers of which 53 were held, so a forward-return test was n=88; the bounded universe backfill closed that (2,786 of 2,792 scored names carry daily closes, 988 of the 1,059 scored on 2026-06-08 span that scan to today), and what remains is TIME — that cohort's 126-day window closes 2026-10-12, so re-weighting on the literature alone would still substitute one unvalidated prior for another. The dependence is instead DISCLOSED to the client on every affected recommendation (see strategy_advisor 1.2.0); nothing about the score changed.
- Captive-finance debt (GM Financial, Ford Credit, Cat Financial, John Deere Financial) is consolidated in every feed line, so an auto or equipment maker with a captive arm scores leverage on the consolidated balance sheet; the Street looks through to the industrial company. Disclosed, not adjusted (2026-09-16: Auto Manufacturers 4 of 4 AVOID).
- Action thresholds live OUTSIDE the label ladder (CRITICAL_SCORE for SELL, PREMIUM_SCORE for the premium screens). They are named rather than bare, but they are judgement calls.
- The headline score is a UNIVERSE percentile and the pillar bars are cohort standing, so the headline is not the sum of its bars (the deliberate consequence of the C5 fix).
- Labels are relative by construction: a fixed fraction of the universe earns each label however good or bad the universe is. The cliffs are the knob.
- Factor weights mirror the legacy scorer's point allocations; the prior is inherited, not derived.
- Below the minimum cohort a sector cannot be z-scored against a real peer group, and the composite is not ranked at all.
- Balance-sheet financials (banks, lenders, broker-dealers, insurers, mortgage REITs) are scored on an UNWEIGHTED capital measure — total assets over tangible common equity. CET1, risk-weighted assets, tier-1 capital and loan-loss provisions are Basel/OSFI Pillar 3 disclosures absent from the vendor feed, so Canadian banks score worse than US regionals partly because insured-mortgage risk weighting is invisible here. No risk-weighted ratio is synthesised from unweighted inputs.
- 1.3.11 (2026-09-17, found by the owner from a Yahoo screenshot, and by the guardrail's first live run): AN UNKNOWN MARKET CAP WAS A ZERO, AND THE FIGURE WAS AVAILABLE ALL ALONG. `info.get('marketCap', 0) or 0` read only quoteSummary, which omits marketCap for real, actively traded companies: on the 2,892-row production scan of 2026-09-17, AutoZone (Deep Score 87, price 2,866.63), TD SYNNEX, ITT, Donaldson, Commercial Metals and MSC Industrial all stored 0. Yahoo's quote endpoint (fast_info) carries the figure for every one of them - AutoZone at 46.7B against the 46.51B on its own Yahoo page - and those same names return no sharesOutstanding either, so shares x price would not have rescued them. `yahoo.market_cap(info, ticker)` is now the one resolver: quoteSummary, then the quote endpoint, and None rather than 0 when neither knows (NaN, inf and non-positive are all rejected as 'not a market cap'). Used by the scanner and by the deep-dive page, whose seven arithmetic sites are guarded - enterprise value, price-to-sales and the market-cap line are now absent rather than computed from a zero, since debt minus cash alone is not a small EV. Real-engine check: all six store a real figure, controls (Apple, Coca-Cola) unmoved, and the two dead tickers are skipped rather than persisted. GUARDRAIL: mcap_sanity keeps flagging a stored non-positive value (nothing writes one now), and a new mcap_unknown_but_priced flags a row we could PRICE but not SIZE - both Yahoo sources failed, so the row cannot be size-tiered or size-ranked. Its close_price read degrades on a schema that predates the column rather than raising, because this guardrail runs when something is already wrong. ALSO CHECKED, and disclosed rather than fixed: (a) the sign_conflict warning RATE rose from 0.6-0.7% on the old universe to 2.6% (74 of 2,892), which is composition not regression - 52 of the 74 carry negative book equity, where a loss over negative equity is positive by arithmetic, and that is the population the expansion admitted; of the 21 with positive equity, six sit within rounding of zero (Air Products' trailing net income really is -47.3M beside a 0.023% ROE) and the rest are a vendor basis difference, netIncomeToCommon being after minority interests while returnOnEquity is not, which is why both Brookfield vehicles appear twice; (b) the scanner's FILTER-metric block still writes 0 for an unmeasurable value - current_ratio 405 of 2,892 rows (14.0%), ebit_to_cap 387 (13.4%), fcf_to_ni 25, altman_z 13 - none of which is ever NULL. Deep Score is NOT exposed: deep_score_features maps a stored 0 to None at the factor boundary for pe, peg and roe ('0 is not a real value'), and current_ratio, altman_z, ebit_to_cap and fcf_to_ni are not factors at all. The harm is confined to the screener's own filters and sorts, where an unmeasurable value reads as the worst one; (c) no reporting-currency column exists at all, so VinFast's net income is stored as -109.8 trillion dong beside USD peers - the deferred native-currency item.
- 1.3.10 (2026-09-17, an adversarial pass over 1.3.9 found two defects IN IT, plus five smaller ones; all closed): (a) THE ABSTAIN DECIDER WAS READING THE WRONG THING. It tested the z-map, and robust_z is None both when a row has no value AND when the COHORT carries fewer than five of them (winsorized_moments' floor), so it could not tell 'this company has no growth history' from 'few of its peers do' - and it abstained on the second. Constructed cohort of 24 Healthcare names with 20 pre-revenue: the four COMMERCIAL names with real growth rates lost their scores, with their own numbers sitting right there; a hard cliff at exactly five peers (4 -> all None, 5 -> 68/76/84/95/100). Latent on the current universe (0 rows affected on the 1,922-row batch, every cohort happened to carry five) and reachable on the population the expansion admits, since _MIN_SECTOR_COHORT is 20 and Healthcare alone contributes 48 clinical-stage names. The decider now reads the row's own factor VALUES; the z-map's None keeps its documented meaning, the too-thin-cohort midpoint. (b) AN UNSCOREABLE ROW WAS STILL SHAPING ITS PEERS. The abstain returned from inside the scoring loop, after the row had already passed through _compute_moments, the per-pillar distribution and the cohort size - so a row that could not be compared was still moving everyone else's percentile (3 of 40 randomised cohorts, and 1.3.9's own test asserted the opposite on fixture luck). Such a row now leaves the cohort entirely, as a no-history row always did. MEASURED consequence, which is the real one: against the pre-change ranking of the same 1,922 rows, 50.5% of scored rows move at all, 9.8% by more than 5 points, 4.0% by more than 10, median 1, p90 5, max 30, and 7.2% change label - and 82 of the 84 names moving 10+ are in Healthcare and Basic Materials, the two sectors whose cohorts had been padded with companies nobody could compare. That is the fix working: real companies are now ranked against real peers. (c) A LAPSED REVENUE TAG INSIDE THE NET-INCOME WINDOW was still served: AGNC files revenue 2016-2019 and net income 2016-2025, so net-income-driven columns brought the lapsed years back INSIDE the window and compute_metrics reported a 2016-2019 revenue CAGR and an 80.6% average margin as current growth and profitability - the NextEra defect re-opened for the in-window case. Revenue is now dropped outright when it is not current. (d) THE PAGER WAS SILENT IN THE ONE CASE IT EXISTS FOR: run_checks itself raising (a corrupt, locked or missing scan DB) skipped main()'s page entirely, and scan_dispatch logs only the exit code; main now pages guardrail_crashed and re-raises. (e) One absent column silenced ALL of section 4 rather than the checks that need it. (f) _knows_latest_revenue selected the newest column by POSITION, which is the newest year in an EDGAR frame and the oldest in a yfinance frame (they are ordered oppositely) - latent, since only the EDGAR frame was asked, and now selected by column value. (g) portfolio_core's book-average Deep Score summed `score or 0`, averaging an unscored holding in as a zero-quality business and feeding it to the AI prompt; it now averages over the scored holdings and states the denominator. The equivalent line in the retired Dash portfolio.py is left alone. Every one of these is gated, each proven red against the code it replaced.
- 2.0.0 (2026-09-18, independent adversarial verification of the whole remediation — five agents, refute-by-default, every high-harm finding re-measured by hand; the engine work held and SEVEN published claims did not, five of them fixed here): (a) A LOW RANK STILL SOLD. Batch 3 gated one rule on absolute_deterioration and the class was reported closed; three paths stayed open. Two paired the percentile with the OWNER'S COST BASIS — a gain or a loss is a fact about the purchase, not the business — and portfolio_db.auto_recommendation asserted 'quality has deteriorated below investable level' from a rank. The gated rule made it worse by standing down for score<35 + gain>25 to avoid duplication, handing the deepest-percentile cases to the one rule with no gate. All three now require an absolute fact and NAME it; the gate is written over the _SELL_RULES registry so a new rule cannot re-open the class. (b) BATCH 4 CREATED THE ZERO IT WAS REMOVING. Making an absent capex line unknown was right, but fcf_consistency counted positive years over total_years, the count of ALL fiscal years, so an unknown year counted as a failure and a company with no measurable year scored 0.0% — the worst value in the field, presented as measured, in a factor worth 7 of 100 points (10 for insurers). It now takes numerator and denominator from one series and abstains when nothing is known. MATERIAL RE-RANK, stated as one: 796 of 2,587 rows scored under both change score by up to 27 points and 96 change label, 282 up and 514 down, and 2 royalty trusts with no measured year stop being scored. RTX and DLTR each had 5 of 11 year columns measured and all 5 positive — 45.5% before, 100% now; LIN 63.6% to 100%. (c) BATCH 12'S PROVENANCE WAS DISCARDED. deep_scored_as was computed and was in no column, no write-back and no consumer, so 127 of 2,589 scored rows published another listing's verdict with no trace — BAM carrying BAM.TO's STRONG BUY 94, GOOGL carrying GOOG's 98. Persisted now, and ranked_batch.leaders_only gives one issuer one slot in the screener and the advisor shortlist; both tickers stay scanned and searchable. (d) THE SUPERSEDED DESCRIPTION WAS LIVE ON THE PUBLIC SITE: compoundwise.io/research served 'percentile-ranked against sector and size peers' and 'deliberately price-blind' from a landing component on the very page Batch 8 corrected, and the in-app Guide, the holdings tooltip and the Android detail screen carried it too. The methodology gate went from 5 surfaces to 10 and stopped requiring the offending line to also say 'deep score' — appending the Guide's exact wording to a gated surface had been proven to leave the suite green. (e) THIS REGISTER published a measured ZERO for the mid-rank percentile change that was 1,632 of 2,589 scores; corrected in the 1.8.0 entry above and gated. Six existing tests asserted the old behaviour and were RE-EXPRESSED, not deleted, each saying why. NOT FIXED HERE and named so it is not mistaken for done: the schema migration still runs on an HTTP request path; winner_engine carries an inline copy of the percentile defect Batch 7 removed; fcf_yield over 2,525 rows bypasses the currency guard; get_usd_cad never raises and returns a hardcoded 1.36; the coverage floor empties every ideas screen after a refused partial scan; and deep_scored_as is not yet on the wire, so no client can say whose answer it is showing.
- 1.9.0 (2026-09-17, Batch 10 — the duplication that is ARITHMETIC, removed; the duplication that is merely CORRELATED, measured and left): the factor-correlation diagnostic built in Batch 9 flagged nine factor pairs at or above |rho| 0.70 on its first run, six of them inside one pillar. They are not one kind of thing, and the distinction decides what may be changed by rule. A factor that is ARITHMETIC on other scored factors in the same pillar counts one measurement twice BY CONSTRUCTION and no market behaviour can separate them: `earnings_growth_premium` is earnings growth MINUS revenue growth and sat in the same Growth pillar as both (+0.82 with earnings growth over 1,098 names), in the default and spread sets. Removed; its 5 points return to the two measurements it was built from (revenue growth 10->12, earnings growth 10->13, pillar unchanged at 25). A factor that merely CO-MOVES is a fact about companies, not a defect — return on equity and return on invested capital correlate +0.81 because profitable firms are profitable, and which of two real signals to keep is a judgement, not a rule. Those are left alone, named, and measured. One case sits between: `operating_margin`, `cycle_avg_op_margin` and `trough_margin` are the last value, the mean and the minimum of ONE margin series, so none is arithmetic on the others and no structural rule reaches them — but the mean and the minimum correlate +0.87 across 218 producers, so the trough was restating the average while holding 6 of the pillar's 30 points. Removed on that evidence, weight to return on equity, the one Profitability factor in the commodity set not drawn from that series. Gate: deep_score_model.DERIVED_FROM names each arithmetic pair and tests/test_deep_score_registry_parity.py fails the build if a derived factor is ever scored beside its inputs again, proven RED by restoring the premium. Measured on the 2026-09-17 snapshot: 166 of 2,892 labels move (default 119, commodity 31, spread 16), no row becomes unscored, and the diagnostic re-run drops from nine flagged pairs to six. STILL OPEN, named not fixed: return on equity vs return on invested capital (+0.81, 22 of 30 default Profitability points), revenue vs earnings growth in the spread (+0.74) and commodity (+0.80) sets, and current vs through-cycle operating margin (+0.73). Each needs a decision about which real signal survives. ALSO IN THIS BATCH, unrelated: the concurrent-replay refresh-token test was flaky at about one run in five. Root cause measured rather than guessed — the server hands every racer the SAME successor jti correctly, but the route re-mints a JWT around it and iat/exp are whole-second, so a burst straddling a second boundary produces different BYTES for the same successor (captured: jti equal, iat one second apart, every other claim identical). A test defect, not a product defect: it asserted byte-equality where its own docstring names the successor as the invariant. Now decodes and compares the jti, plus asserts each racer can USE what it was handed; 0 failures in 40 consecutive runs, and still RED when the route is made to hand back different successors.
- 1.8.0 (2026-09-17, Batches 6-9 — beta priced once, the percentile made unbiased, the published description reconciled, and the forward read made able to grade): BETA LEFT ALL SEVEN FACTOR SETS. It was a Moat & Quality factor, and that pillar feeds quality_score_from_breakdown, which sets the +/-100bp step in cost_of_equity beside the CAPM term that already prices beta — so a low-beta name earned a quality credit that cut its own discount rate and raised its own fair value. Its 4 points went to EARNINGS VARIABILITY, an idiosyncratic measure with no reference to the market; the pillar stays TWO-factor in every set because folding the weight into a single remaining factor collapses the pillar whenever that factor is unmeasurable, which under the abstention rule unscores the row (measured: 101 rows, including a producer at 99 and a REIT at 100 whose margins are perfectly measurable). Effect on the 2026-09-17 snapshot: 613 of 2,892 labels move and 96 rows become unscored, 80 already AVOID, concentrated in Basic Materials 34 / Healthcare 25 / Technology 19 — beta had been propping up a BUSINESS-quality pillar for pre-production miners and clinical-stage biotech whose business quality cannot be measured at all. THE PERCENTILE IS MID-RANK AND SELF-EXCLUSIVE: counting at-or-below over the whole field including x made a cohort's own mean 100(n+1)/2n rather than 50 and gave a tied block its TOP rank. Measured effect, CORRECTED 2026-09-17 after independent verification: isolating this change on the HEAD engine over the same 2026-09-17 snapshot moves 1,632 of 2,589 scores, by up to 7 points, and changes 109 labels. The entry first published here claimed no score moved at all, reasoning from the 100/(2n) self-inclusion bias over cohorts of 27 to 375. That argument is correct and covers only HALF the change: it drops the TIED-BLOCK half, and a tied block is most of the field on a discrete factor such as a count of positive free-cash-flow years, so those names move by whole points. The fix itself stands — a field's own mean is now its midpoint and a tied block shares one rank — but it is a material re-rank, not a silent correctness tidy, and it is gated by tests/parity/test_model_register.py so the zero cannot come back. THE PUBLISHED DESCRIPTION was reconciled across eight surfaces (composite formula, the removed free-midpoint rule, the quadrant rule that gated on a P/E z and was wrong for every class, the sector-and-size claim on the headline in tooltips/glossary/research/report prompt, the pillar-summing score simulator, the never-rendered confidence flag) and gated by tests/parity/test_methodology_claims.py, which reads the real constants. THE FORWARD READ CAN NOW GRADE: forward_return separated IMMATURE (window not closed) from DELISTED (stopped trading) — they had been one branch and are opposite facts — and the advisor track record, which had read its (ret, status) tuple as a scalar and raised on its first graded pick since 464b130e, produces a number for the first time; the excess benchmark now keys on the PEER GROUP the score was computed in rather than the raw sector, so a bank is measured against spread lenders and not against a pool holding asset managers and exchanges. NEW: scripts/deep_score_factor_diagnostics.py measures what the audit could only read — within-class pairwise Spearman between every factor pair, coverage per factor, and the Deep-Score-vs-Inflection correlation that inflection.py has always ASSERTED is orthogonal. First run on the 2026-09-17 snapshot: the orthogonality claim HOLDS at rho +0.126 over 2,403 rows, and nine factor pairs sit at or above |rho| 0.70, six of them inside one pillar — commodity through-cycle margin vs trough margin +0.87, default earnings growth vs the earnings-outgrowing-revenue premium +0.82, ROE vs ROIC +0.81. Those are named for a later batch, not fixed here.
- 1.7.0 (2026-09-17, Batch 5 — each per-class factor now computes what its NAME says): nine units. (1) The regulated-utility credit factor published as FFO / debt computes OPERATING CASH FLOW / debt — the rating agencies' measure is cash earnings BEFORE the working-capital swing operating cash flow contains, so a receivables timing item moved a credit factor by hundreds of basis points and a reader comparing it to an S&P threshold compared two different measures. The feed carries no cash-earnings line for utilities, so it is RENAMED rather than re-derived. (2) The commodity valuation factor divided an EQUITY price by UNLEVERED after-tax operating earnings and published it as Price / through-cycle earnings, making leverage invisible: a producer with billions of net debt scored as cheap as a debt-free peer in the one class where leverage decides who survives the trough. Both sides now include debt — EV / through-cycle operating earnings. (3) The cycle factors read ONE window. Depth varies by SOURCE, not by company: 189 of the 259 scored producers carry exactly 4 years of operating margin (vendor frames) and 39 carry 10 (EDGAR), so a 4-year through-cycle average was z-scored against a 10-year one and the factor partly ranked data depth. Capped at 4, the depth almost every name has. (4) The software Growth pillar scored revenue growth TWICE — standalone and inside the Rule of 40 — while the Rule's other half, FCF margin, is already the SBC-adjusted FCF margin factor in Profitability, so a fast-growing cash-generative name earned the same two facts three times and 80% of a 25-point pillar was one number. Standalone revenue growth is dropped and its 10 points go to the Rule of 40, the composite the industry actually quotes; the pillar weight is unchanged. (5) Nareit FFO removes the gain on a property sale, but 27 of the 118 scored equity REITs carried no gains line under the four mapped tags, so a disposition read as FFO growth — the heaviest factor in the class — and overstated ffo_yield, ffo_roe and ffo_margin with it; six more disposition concepts added. (6) FFO margin divided last-ANNUAL FFO by _latest(revenue), which can be the TTM point compute_metrics appends — a period mismatch that reads as a margin collapse for a REIT that just closed an acquisition; the two are now taken from the same year. (7) Gross margin survives a lapsed GrossProfit tag: 43 of 191 scored software names carry no gross-profit line while still reporting revenue and cost of revenue, so cost_of_revenue_data is now extracted, persisted and used as the fallback. (8) The insurer loss ratio is named CLAIMS / TOTAL REVENUE (LOSS-RATIO PROXY) wherever a reader sees it — the denominator includes investment income because the feed has no premiums-earned line, so for an equity-heavy insurer it partly tracks the market rather than underwriting; disclosed in the NAME, not only in a code comment. (9) A mortgage REIT's vendor revenue is interest income GROSS of the funding cost where a bank's is net interest income plus fees, so net margin, revenue growth and the earnings-outgrowing-revenue premium are not the same quantity and were z-scored against 208 banks as though they were; those three are unknown for a mortgage REIT and the pillar re-bases. Earnings growth is KEPT deliberately — it reads net income, which means the same for both, and blanking it too empties the Growth pillar and makes every mortgage REIT unscored, a larger error than the one being fixed. Measured on the 2026-09-17 snapshot, 127 of 2,892 labels move (commodity 55, software 29, default 27, spread 8, equity REIT 7, insurance 1); the mortgage REITs re-rank hardest (NLY 67 to 42, EFC 56 to 33) because the revenue-line factors had been flattering them, and BXMT becomes unscored because its remaining Growth factor is unmeasurable. Gates in tests/test_deep_score_class_factors.py, one per unit, each proven RED against the defect it names. The class-invariant suite now reads its per-class signature factor from the REGISTRY rather than restating the name, because a rename had silently un-tested the insurance class.
- 1.6.0 (2026-09-17, Batch 4 — an absent capital-expenditure line was a zero, so free cash flow was operating cash flow): fundamentals_metrics computed `ocf + (capex or 0.0)`, which silently republished operating cash flow AS free cash flow for any filer whose capex line the EDGAR tag map did not carry. Kosmos Energy stored +765M, +678M and +134M for three years it burned -167M, -256M and -180M. Measured on the 2026-09-17 scan, 285 US rows in the operating classes (the spread and insurance sets score no FCF, so they are unaffected) had latest-year FCF exactly equal to OCF with capex absent: 152 default, 80 equity REIT, 29 commodity, 13 regulated utility, 11 software. Three changes. (1) FREE CASH FLOW NEEDS BOTH HALVES — an absent capex is unknown capex and FCF is None, so the 58 of 108 probed filers still without a capex line now report nothing rather than something wrong. (2) THE TAG MAP IS WIDER, measured rather than guessed: probing SEC companyfacts for 115 of the affected filers found 61 already carrying a mapped tag (missing the YEAR, not the concept) and the rest split across industry concepts, so PaymentsToAcquireOtherProductiveAssets, ...OtherPropertyPlantAndEquipment, PaymentsForCapitalImprovements, the four oil-and-gas development concepts, machinery, buildings, mining and capitalised software were added — 50 of 108 recover a real latest-year capex. DELIBERATELY EXCLUDED, because they would change what free cash flow MEANS rather than complete it: PaymentsToAcquireRealEstate (30 filers) and ...CommercialRealEstate (11) are a REIT buying whole income properties, growth acquisitions rather than maintenance of assets already owned; PaymentsToAcquireIntangibleAssets (21) is not PP&E; PaymentsForProceedsFromProductiveAssets (7) is a net figure whose sign flips with disposals. (3) CAPEX IS THE PER-YEAR MAXIMUM ACROSS ITS TAGS, not _pick's newest-tag rule — the same treatment Net PPE already has, and necessary because a widened list lets a minor line that is still filed shadow a principal programme that lapsed (Murphy Oil reads $8-36M against a ~$1B programme under the OTHER- sibling). A filed total is never smaller than one of its components, so the max is a floor on capital spending, never a ceiling. Also recorded, reading nothing yet: `debt_basis` per row — which of the six Long Term Debt concepts (or yfinance's lease-inclusive total for .TO names) a row's leverage came from, so the mixed-basis cohort finding can be MEASURED before it is fixed. Gates in tests/test_fundamentals_metrics_unknowns.py (FCF is None when capex is, and a vendor FCF line still wins) and tests/test_sec_edgar_facts.py (the industry concept is read; a minor concept cannot shadow the principal one in a year both are filed; debt_basis is recorded), each proven RED by restoring the defect. RESIDUE, named rather than closed: 58 of the 108 probed filers still file no capex concept this map reads, and _pick's one-winning-tag rule still under-counts a filer that splits capital spending across two concepts.
- 1.5.0 (2026-09-17, the first batch of the Deep Score audit remediation): A PORTFOLIO IS NOT A PEER. Closed-end funds and business development companies were classified as spread lenders and ranked inside the bank cohort: on the 2026-09-17 production scan, 36 of the 250-name cohort, six of them publishing as STRONG BUY (Hercules Capital 100, Main Street Capital 98, PIMCO Corporate & Income 98, PIMCO Dynamic Income 95) against a median of 49 for the 208 real lenders — and every real bank's leverage z-score was computed in a pool that held them, a fund at 1.72x assets-to-tangible-equity beside a bank at 11.64x. A fund's investment income IS interest income, so the 10%-of-revenue test can never separate the two; the BALANCE SHEET can. business_class gains an INVESTMENT_COMPANY class decided by two POSITIVE signals, never by an absence: (1) the filer's own facts — a BDC files us-gaap InvestmentOwnedAtFairValue AND NetInvestmentIncome, present for 15 of 15 BDCs and neither present for any of 12 banks and brokers, now recorded by sec_edgar and persisted by the scanner as is_investment_company; and (2) the balance sheet, for the closed-end funds that file N-CSR and reach no companyfacts at all — a name presenting as a lender by income while carrying under 3.0x assets per dollar of tangible equity (INVESTMENT_COMPANY_LEVERAGE_MAX). "EDGAR returned nothing" is deliberately NOT used as a signal: a 404, a timeout and a fund are indistinguishable. Such a row is UNSCORED by deep_score_engine, partitioned out above the cohort floor so it shapes no peer group's moments, count or size split, and _score_group now RAISES if one reaches it. The fair value moves with the score rather than after it: valuation/class_policy excludes every one of the 21 models for the class, so the ensemble returns None at NOISE — suppressing only the score would have left PIMCO Dynamic Income carrying a DCF, and with it a margin of safety, a position size and an advisor recommendation. Measured on the 2026-09-17 snapshot: 35 rows reclassify (31 known funds with none missed, plus Clairvest, Strive, Hyperliquid Strategies and Lufax), the spread cohort falls 250 to 215 leaving exactly the three custody banks and brokers the taxonomy intends, 18 scored rows become unscored, and total label churn is 2.9% of 2,892 — of which 31 are names sitting exactly on the 85 cliff moving to 84 as the quota reallocates. ONE THING THE AUDIT PROPOSED WAS NOT DONE, on evidence: bounding the net-interest share to 10-100% and treating an out-of-band value as absent reclassifies exactly three rows on that scan and all three are real consumer lenders wrongly demoted (Synchrony 123.3% then scoring 90, Bread Financial 105.5%, Atlanticus 273.7%), while rescuing no fund. A share above 100% is an unreliable magnitude and still unambiguous evidence of interest-income orientation, so the taxonomy reads it as the pass it is; the scanner logs it and scan_integrity pages on it as the DATA defect it is (45 of 2,892 rows, 6 of them genuine banks). Gates: tests/test_business_class.py (a 14-row balance-sheet corpus plus the hand-reviewed data/class_review/ file, 169 rows, re-read on every run), tests/test_deep_score_engine_golden.py (the fund is unscored AND no bank's score, label, cohort size or tier moves because it was in the batch; _score_group refuses it), tests/valuation/test_class_policy.py (no model is included and the blend returns None), tests/test_sec_edgar_facts.py (both tags required — InvestmentOwnedAtCost alone false-positives on Eastern Bankshares). Each was proven RED by injecting the defect it names. The pre-flight is now scripts/deep_score_prod_diff.py, which runs the engine over a production snapshot and exits non-zero on a named gate, replacing the eyeball review the class rollout rested on.
- 1.4.0 (2026-09-17, the rollout ENDS): the per-class programme is no longer switchable. DEEP_SCORE_CLASSES (and its 2026-09-15 alias DEEP_SCORE_BALANCE_SHEET_BASIS) is RETIRED — the default is flipped in code and the branch deleted, so business_class.business_class() is what scoring, valuation, the peer groups and the persisted deep_class column all read, with no environment gate. Retired on evidence, not on a date: both deploy targets served commit 666ba5bc with all six classes enabled, and the production fundamentals table carried 320 non-default rows of the 1,037 stamped with a class, every class represented. Two behaviours rode the same switch and are now permanent with it: the balance-sheet basis for spread lenders and insurers, and the DYNAMIC peer medians for P/E, P/B and P/S in data/resolver (static Damodaran tables remain the fallback when a caller has no peers to measure). DEFAULT is unchanged in meaning — it is the taxonomy answer for an ordinary operating company, never an off state. Gates: tests/test_business_class.py fails the build if the retired token reappears anywhere under src/ or if the deleted symbols come back, and tests/parity/test_flag_inventory.py fails if the switch is re-declared. The on/off laws that pinned the rollout are re-expressed against a DECLASSED universe (the two taxonomy columns nulled), which is the only honest way left to reproduce the pre-class engine.
- 1.3.9 (2026-09-17, the three residuals 1.3.8 disclosed, closed): (1) THE READER no longer needs a revenue tag. Revenue drives the columns while it is CURRENT; where a filer tags none at all, or none since years ago, net income drives them and revenue is left unknown - so a clinical-stage biotech, a pre-production miner and an agency mortgage REIT are read rather than discarded. The annual report may also be a 20-F or 40-F: a foreign private issuer reporting under US GAAP files the same us-gaap facts, and a 10-K-only form filter had been reading JD (2014-2025), Melco (2008-2025), Radware and Nordic American as having no financial history at all. Measured: of the 108 rows the sweep found with NO history, 101 now have one (7 remain, all closed-end funds and shells); on the 200-name census the reader parses 143 (72%) against 131 (66%), and 137 (68%) are scanned from EDGAR. GUARD: EDGAR is preferred for depth, not unconditionally - where its revenue is not current and yfinance's is, the shallower frames win, so Meritage Homes (17 years of net income, revenue only under a company-specific extension tag) keeps four real revenue years instead of ten blank ones. Controls unmoved: KO, JPM, AAPL, EQR read 10 EDGAR years. Still on yfinance and now the whole of the residual: filers with no us-gaap facts (IFRS 20-F/40-F: AstraZeneca, BCE, Enbridge, Cenovus, Diageo, Lloyds), plus BDCs and closed-end funds. (2) A MINIMUM-FACTORS RULE, taken as full pillar coverage. A pillar with no measurable factor was scored at the neutral midpoint - half its points, free, indistinguishable from a company measured and found average - so the row is now unscored, exactly as one with no history is. Measured on a realistic 1,922-row batch rather than a throwaway one: 104 rows were scored on four pillars, and removing the free midpoint moved 89 of them DOWN by a median of 4 while pushing the strongest UP (Airgain 82 BUY -> 90 STRONG BUY), because the unmeasurable pillar had been dragging a well-ranked name toward the middle - wrong in both directions, which is why re-basing over the pillars that exist was measured and then rejected in favour of abstaining. It costs nothing on real companies: 0 of the 1,008 hand-curated names is partial, while the 104 are 48 clinical-stage biotechs and 41 pre-production miners whose missing pillar is Growth (84) or Profitability (20), never both and never any other. 133 of 1,922 rows (6.9%) are now unscored, and every scored row carries all five pillars. Scored rows move by at most 3 points, from the batch distribution changing when the unscoreable rows leave it, not from any input changing. CORRECTION to 1.3.8: the "Abivax 83/BUY" quoted there was measured in a 36-name throwaway batch and was never a universe result; on the real batch Abivax scored 66/HOLD before this change and is unscored after it. (3) THE GUARDRAIL PAGES. scan_integrity.main() now sends the report to the owner's Telegram on any error, once per day per problem SET (the scan runs four times a day; a new problem pages immediately, and tomorrow's scan pages again). It posts with requests rather than through alert_engine, which imports redis and yfinance at module scope - the guardrail that runs because the scanner misbehaved must not be takeable down by the scanner's own dependencies. An unconfigured token, a failed send and an unwritable marker each print loudly; none is silent. Section 4 also gains partial_framework_scored, the new pillar rule checked at the data edge where a second writer would show up. Behaviour change disclosed: Discovery approve answers 422 for a name the framework cannot measure, as it already did for one with no history.
- 1.3.8 (2026-09-17, the verification sweep over the 1,922 names the 2026-09-16 ledger never saw, run on the real engine then rank-passed; four ways an unknown became a number, each fixed at its choke point and re-verified on the names that showed it): (a) a fiscal year existed only if revenue was positive - Yahoo serves a pre-revenue filer's revenue as a literal 0.0 beside a full net-income column, so Viking, Biohaven, QuantumScape, Oklo and NexGen (20 of 20 sampled, 4-5 years each) had NO history: 108 of 1,922 swept rows (5.6%), 0 of 1,008 in the curated universe; a year now exists when revenue > 0 OR net income is known, revenue None-preserving (a reported 0.0 stays a measurement); (b) partnership equity read as 0.0 - Yahoo files 'Stockholders Equity' as exactly 0.0 for every LP (BIP, BEP, BBUC, EPD, MPLX, ET) with the capital under 'Common Stock Equity', and the EDGAR reader had no PartnersCapital concept, so equity was [0.0, 0.0, 0.0, 0.0] beside $69B of debt (6 swept rows; BIP-UN.TO / BEP-UN.TO before) and crit_bv_growing was a FAILED test rather than an unknown one; the fallback fires on the artefact (a literal 0.0) or an absent row only, never on a single NaN year (a corporation with preferred equity would read a mixed basis), and EDGAR admits PartnersCapital only where the stockholders' line is absent or lapsed (EPD/MPLX/ET now read 10 years; a stray newer partners fact cannot displace a corporation's history under _pick's newest-wins rule); (c) the TTM free-cash-flow slot defaulted to 0.0 when nothing knew it ([None, None, None, 0.0] on 14 closed-end funds and royalty trusts) - now None; (d) a row with no history was scored at the neutral midpoint (every pillar 50%, every note n/a, then universe-ranked: the 108 landed on median 56, 76 HOLD, confidence 'medium', which the advisor's requires('score') gate reads as measured) - such a row is now UNSCORED (deep_score NULL) and kept out of its cohort's moments and count (a 19-name sector plus one unscored row is a 19-name cohort and scores on the universe rung: measured 0 to -7 on that fixture, deliberate), the harvest sentinel no longer persists deep_score 0 between the harvest and the rank pass, and the deep-dive page reads the last KNOWN point of a None-bearing history (it raised a TypeError on the new shapes and /stocks/<t>/live rendered the crash as 'no data'). Proven on the pre-fix sweep data: exactly the 108 zero-history rows come out NULL, 1,814 keep scores. Behaviour change disclosed: Discovery approve now answers 422 for a name with no financial history instead of approving it at score 0. Standing gate: scan_integrity section 4 runs these checks on every scan as batch RATES (unknown_scored, zero_equity_history, zero_debt_history_rate, no_history_rate, sector_unknown_rate), and tests/invariants/test_series_builders_never_default_to_zero.py fails CI if any per-year history in compute_metrics spells a missing cell as a number. Residuals: sec_edgar still returns None for a pre-revenue 10-K filer (revenue years drive its columns), so those names read 4-5 Yahoo years rather than 10; a pre-revenue name now scores on the few factors it has (Abivax 83/BUY at confidence low) - a minimum-factors abstain is a product decision not taken; scan_integrity's verdict reaches only the dispatcher log. EDGAR census n=200 per population: never-swept US names fall back to Yahoo 69 of 200 vs 5 of 200 on the ledger's names, 4 of those 5 dead tickers still in the seed.
- 1.3.7 (2026-09-16, the two follow-ups from 1.3.6 closed): (a) the SEC ticker map falls back to ticker.txt for filers company_tickers.json lacks - 8 of the 946 US names in the universe (AvalonBay, Equity Residential, Electronic Arts, Webster Financial, Chart, National Storage, Taylor Morrison, Bed Bath) were served the 4-year yfinance history for want of a CIK, and AvalonBay's Yahoo counts carry a spurious 2.793 factor; (b) a revenue year filed only under a lapsed tag fills a gap INSIDE the live span (NVIDIA's FY2019 exists only under the ASC 606 tag it used for three years; the live-tag rule had dropped it from the middle of the series) but can still not extend the span with ancient years; (c) the thousands-vs-units share snap runs AFTER split restatement, not before - on raw counts a 40x cumulative split (NVIDIA 2016 at 539M against 24,304M today; Chipotle's 50:1) read as a unit error and 2016-2020 were served as unknown; (d) ASC 842 lessors' lease income (OperatingLeaseLeaseIncome) is a revenue line - Equity Residential's `Revenues` lapsed in 2019 and the filer fell back to yfinance for want of a live tag - and a retired tag extends the live span BACKWARDS where its years are contiguous with it (tag succession), while a retired tag with a gap before the span still cannot union its years in.
- 1.3.6 (2026-09-16, prod rescan review of 1.3.5): two more layers under the filing-vintage rule. (a) A later 10-K carries the prior year-end share count UNRESTATED in its equity roll-forward beside a restated balance sheet (Edwards' FY2020 10-K: 2019 = 209.1M against its own 636.7M weighted average; Cognex's FY2017 10-K: 2016 = 84.9M), so 'filed after the split' never meant 'restated': each year's count now comes from the year's OWN (earliest) 10-K with that filing date as its basis, and a year-end count that disagrees with the same filing's weighted average by more than 1.5x is replaced by the average (ASC 260 requires it restated; Cognex reported 85.94M at 2017-12-31 beside a 179.55M average). (b) yfinance's split series carries spin-off and special-dividend PRICE factors that never changed a share count (EQT 1.837, AvalonBay 2.793, Crane NXT 2.879, Dell 1.806): a split is applied only where the as-reported counts confirm it within 15% of the ratio (a factor within 15% of 1 is never a split: 33 of the 240 vendor factors across the US universe are 1.02-1.18 dividend and spin-off adjustments a flat series would 'confirm'), or where a clean small-integer ratio (2, 3, 0.2, 5/3) is not contradicted by a flat series - AMC's 1:10 reverse split coincided with a unit conversion (observed 0.32) and is still applied; AvalonBay's 2.793 against a flat 142.8M today is not. One rule now covers EDGAR (vintage = own filing date) and yfinance (vintage = year-end) series. Real engine 2026-09-16: Edwards 634.8-580.7M smooth (was a 3x step at 2019), Cognex 171.9-167.0M (was 85.9 | 179.6), EQT and Ovintiv untouched by their factors, RLI/TPL/Apple as in 1.3.5.
- 1.3.5 (2026-09-16, prod scan review of 1.3.4's decision 3): a share count is restated by its FILING vintage, not by a blanket split multiplier and not by a jump test. EDGAR keeps the latest filing per year, and comparatives filed after a split are already in post-split shares, so multiplying every earlier year by the split history counted each split twice (RLI's 2:1 read 4x, Texas Pacific Land's two 3:1 read 9x); applying a split only where the count jumps by its ratio then missed the case where one series mixes vintages (the oldest years from 10-Ks filed before the split, the rest from the 10-K filed after it - a 2x step left inside RLI, 3x inside TPL, 4x inside Apple). Now each EDGAR count carries the filing date of the fact it came from (sec_edgar `shares_filed`) and is multiplied by every split dated after that filing; a yfinance series, which carries no vintage, is restated only where the count jumps by the split ratio. Real engine 2026-09-16: RLI 87.9-91.9M against 91.8M today, TPL 70.1-68.9M against 69.0M, Apple 2017 restated 20,505M.
- 1.3.4 (2026-09-16, the four decisions from the universe verification, in priority order): (1) the insurance set's leverage factor is FINANCIAL leverage — debt / (debt + shareholders' equity), the S&P / Moody's debt-to-capital measure — not assets / tangible equity, which counted policyholder liabilities as leverage and pinned every life insurer at z = -3 on 8 points (MET, PRU, MFC, GWO at 25-80x beside P&C at 3-6x; Insurance - Life median 24, 64% AVOID); book equity, not tangible, because a life insurer's intangibles line carries DAC/VOBA. The spread set keeps assets / tangible equity: a bank's balance-sheet leverage IS its risk. (2) ETFs (etf_scoring.py) are judged within an ASSET CLASS inferred from the vendor category or the fund's name (Yahoo has no category for any TSX listing: 138 of 140 Canadian funds) — the performance ladder moves for fixed income, cash and balanced funds while the points per rung do not; and a return window a fund is too young to have is unknown, not 0 (CASH.TO, CBIL.TO, TCSH.TO scored 0 of 35 for a history they cannot have), the score re-based on the pillars that exist. Dry run on the prod copy: fixed income median 57 -> 65, cash 48 -> 69, equity unchanged. (3) Historical share counts are restated in today's shares from yfinance split history, so a per-share CAGR never reads a split as dilution (RLI 2:1, TPL 3:1, AAPL 4:1). (4) Captive-finance debt (GM Financial, Ford Credit, Cat Financial, John Deere Financial) cannot be separated from the parent's in any feed line, so auto and equipment makers with captive arms score their leverage on the consolidated balance sheet — disclosed, not adjusted.
- 1.3.3 (2026-09-16, verified against Street consensus): the utility class is the WHOLE Utilities sector - independent power producers and renewables included (the Street scores a merchant on EV/EBITDA, FCF yield and hedged cash flow, i.e. the utility set's asset-base growth, FFO/debt, coverage and payout; as commodity producers VST scored 0 and CEG 14 against Strong Buy / Buy consensus), 'Rate-base growth' renamed 'Asset-base growth (net PPE)'; the commodity set now leads with CURRENT cash generation (current operating margin 10, through-cycle 8, trough 6, ROE 6; net debt/EBITDA 8, FCF consistency 7, coverage 5; current FCF yield 8, normalised P/E 4, tangible book yield 3) - the valuation layer already normalises the cycle and doing it twice erased the current regime (EQT 33 with 24 of 25 analysts at Buy). EDGAR substrate, each verified on the real companyfacts of the named filer: EBIT = pretax + interest for filers with no OperatingIncomeLoss (9 of 38 US commodity filers had margin factors of None; all 9 derive now), pretax bottom-up as net income + tax where no consolidated pretax tag is filed (CNX, RRC), interest from InterestAndDebtExpense (OXY) and, per year, from the net non-operating interest line in expense years only (NEM), and Net PPE as the per-year MAX across its tags plus the post-ASU-2016 combined tag - one winning tag read AEP's plant as [73.3, 0.6, 0.7, 0.7]B (asset-base growth -78%, BUY -> AVOID), ETR and SR at ~1% of their plant, and SO/FE/PCG/SRE with multi-year gaps. NJR files only gross plant-in-service and stays unknown. Universe sweep (SEC bulk companyfacts, every US name): revenue was the ASC 606 SLICE for 58 names (17 banks, 14 REITs, 5 insurers; ADM 25.0B vs 85.5B, MET 2.2B vs 77.1B, COF 5.9B vs 39.1B) - now the per-year max across live revenue tags, with net interest income + noninterest income (or revenues net of interest expense) as the bank definition; the insurer loss ratio double-counted one line filed under two names (PGR 126% -> 63%, HIG 148% -> 74%); a revenue tag that lapsed (NextEra: 2012) had a 13-year gap served as one year of growth - histories not reaching the last two fiscal years are now unusable (yfinance fallback); evidence-based tag additions took unusable filers 59 -> 17, EBIT-None 204 -> 64, D&A-None 134 -> 24, share-count-None 103 -> 26; and the metrics builder wrote 0.0 for an UNKNOWN year (480 of 1,007 rows carried a 0.0 debt history, 94 equity, 60 net income) - now None, with every consumer guarded. Residuals disclosed: 17 IFRS/extension-taxonomy filers on the 4-year yfinance fallback (TSM, ASML, PBR, XOM, APA), 13 of 54 US equity REITs with no interest line under any standard tag (operating margin unknown), two true stock splits in per-share classes (RLI 2:1, TPL 3:1) not yet split-adjusted. These populate on the next scan.
- 1.3.2 (2026-09-16, owner review of the enabled dry run): the commodity set gains a CURRENT operating-margin factor (6 of the Profitability pillar's 30; through-cycle 10, multi-year ROE 8, trough 6) so a producer at peak margins is credited for today's cost position, not read as if at trough; independent power producers and renewables move from the thin Utilities residual (where they fell to the universe pass on the default set — CEG 78 -> 20) into the commodity class as power-price takers.
- 1.3.1 (2026-09-16, T4): the readers follow the class. The three portfolio screens no longer exclude every bank by `debt_to_equity < 1.0` (a spread lender or insurer passes on its class; the D/E clause applies to the rest); signal_engine's dangerous-leverage SELL (D/E > 3 and score < 45) is off for spread lenders and insurers, whose D/E is structural; four checklist criteria (FCF >= NI, FCF growing, FCF quality, improving D/E) are computed for banks and insurers as profitable-every-year, book-value-per-share growing, ROTE >= 10% and tangible equity >= 5% of assets, and the served labels say so; the live 'Quality Compounder' badge no longer demands D/E < 1 of a bank. The class switches (DEEP_SCORE_CLASSES=all, VALUATION_DISCOUNT_RATE=2026) are deploy-environment settings, not code defaults: OFF in code is byte-identical to before, and the owner turns them off by editing the deploy env. Known: the advisor's per-ticker map reads deep_class, which is NULL until the deep-score pass has run on the enabled basis.
- 1.3.0 (2026-09-16): per-class FACTOR SETS (deep_score_model.FACTOR_SETS). Six classes carry a full five-pillar set with the same pillar maxima, so the composite, label cliffs, quality score and every client are unchanged; each class is ranked only inside a cohort scored on the same set (one set per cohort, decided from the peer group). Standards behind each set: Damodaran (financial service firms: ROTE, equity models); MSCI Quality (earnings variability); Nareit FFO (NI + D&A + impairment - gains on sale); the regulated-utility convention (rate-base growth, FFO/debt, earnings payout); Damodaran on cyclical/commodity firms (through-cycle and trough margins, normalised P/E); the SaaS Rule of 40 and SBC-adjusted cash margin; Novy-Marx (gross margin). Limitations by class: insurer loss ratio is a proxy over total revenue (no premiums-earned line); REIT FFO carries no maintenance-capex (AFFO) deduction; utility rate base is net PPE, not the regulatory rate base, and allowed ROE is unavailable; commodity normalisation uses the stored 4-10 years, not a full cycle for every name; software has no net revenue retention. Switched per class by DEEP_SCORE_CLASSES, default OFF until each class's prod dry run is reviewed.
- 1.2.2 (2026-09-16): business_class.py is the single taxonomy (7 classes: default, spread, insurance, equity_reit, regulated_utility, commodity, software), switched per class by DEEP_SCORE_CLASSES (default OFF; the 09-15 flag is an alias for spread,insurance); the scanner now stores seven per-year arrays (D&A, OCF, capex, SBC, operating gains adjustment, impairment, net PPE), an insurer loss-ratio PROXY over total revenue (the feed has no premiums-earned line) and the class each row was scored as (deep_class). No score moved: the arrays feed factor sets that land in 1.3.0.
- 1.2.1 (2026-09-15, same day): the insurance test is industry 'Insurance -' so the 8 Insurance Brokers (fee businesses) stay on the standard basis; the book-value-per-share factor now survives the TTM pad every shares_data series carries (it was None for 108 of 112 balance-sheet names on the first prod dry run) and compounds over the index span of the first and last known years; EDGAR now supplies a share count for US filers.
- The stored fcf_yield column is still the vendor's free-cash-flow yield for every name (TD.TO: -265%); only the SCORE substitutes tangible book yield for balance-sheet financials. The insurer loss ratio (policyholder benefits + loss adjustment expense over total revenue, a proxy — the feed has no premiums-earned line) is a factor of the insurance set since 1.3.0.
- Fed by a single vendor with no point-in-time archive (audit C7). Measured 2026-09-10: between two scan dates a week apart ALL 1,001 tickers moved on avg_roic — which was our own C4 fix, not the vendor, and nothing in the data said so. Scan dates now carry the git SHA that produced them, which makes code-vs-data attributable; reproducing a vendor figure as it stood on a past date remains impossible.
Entry points & executable coverage
Winner Odds
winner · owner vishalsinha · version 1.16.1 · last reviewed 2026-10-06 · winner_engine.py
Monitoring
WIRED AND RUNNING verified against prod; earliest 1yr cohort matures ~2027-07-21
Metric: rank-IC, top-vs-bottom decile forward spread, winner hit rate. Grader: scripts/winner_track_record.py. Log: winner_predictions.
Population test
certified 2026-10-06 scripts/backtest_winner_population.py · re-run on the corrected history (1.16.1): rank-IC 0.234, tercile spread +0.203, t 4.52, deflated Sharpe 0.9997 over N=17, held-out halves 0.228 / 0.226 -> rank certified
Scope: the rank (ordering) over US 10-K filers FY2009-2022; not a probability (CALIBRATED False); TSX listings not tested
Validation
self-gating + invariant The rank: scripts/backtest_winner_population.py over every US 10-K filer FY2009-2022, judged by the one certify() rule (scripts/_winner_backtest_stats.py) — certified at 1.12.0 / 1.16.0. The probability: CALIBRATED=False until scripts/winner_calibration.py passes on held-out blocks (it failed 2026-09-28), and tests/invariants/test_remaining_entry_point_invariants.py ENFORCES that gate: no calibration parameter may be set while it is closed, and thin evidence must abstain rather than be ranked
Assumptions
Limitations
- 1.16.1 (2026-10-06, MODEL UNCHANGED; THE CERTIFICATION RE-READ ON A CORRECTED HISTORY - STILL CERTIFIED): two defects in the shared population harness were fixed for the Deep Score's own test, and this certification was re-run on the repaired history as a recorded trial in its own log. (1) The SIC map tested 2800-2899 before 2830-2836, so drug makers were peered as Basic Materials / Specialty Chemicals (1,381 filers in the pre-repair harvest), and listed 3700-3799 under both Consumer Cyclical and Industrials, so aircraft, ship, rail-equipment and missile makers were Specialty Retail (79 filers). (2) The 2023q1 data set had never loaded (the SEC link 404'd on 2026-09-27), so 3,222 FY2022 rows carried the next year's filing date - a year-late entry - and most of FY2022 was an open window. RE-RUN: held-out halves (winner_factor_attribution.py) 0.228 and 0.226 (were 0.250 / 0.219); rank-IC 0.234 (was 0.244), tercile spread +0.203 (was +0.244), L/S Sharpe 2.02, t 4.52 (was 4.71), deflated Sharpe 0.9997 over N=17 recorded trials (was 0.966 over 15; the two IC-weighted schemes re-fitted to new weights and count as two new trials) -> RANK CERTIFIED. Funnel: 65,318 as-of rows -> 24,037 ranked and voting -> 15,057 graded (14,346 priced, 711 bankrupt). Measured >= +150% base rate: 0.071 to 0.240 across the exit bounds; BASE_RATE 0.083 kept, inside them. NOT RE-DERIVED here: the five-year quartile figures and the calibration read, so every surface that quotes them dates them to the 2026-09-28 history. The documents that quoted the 0.244 / 4.71 certification figures were updated to these in the same release (5f63ef2b).
- 1.16.0 (2026-09-28, THE CERTIFIED FIGURES WERE MEASURED ON A HISTORY THAT READ SOME DEBT AS ZERO; CORRECTED, STILL CERTIFIED; ONE PARSER FIX SHIPPED, ONE DECLINED): the Piotroski coverage funnel on the first scan after the gross-profit and zero-borrowing rules (1,174 of 1,790 US tiered names carry an F-score; 369 lack a cost-of-goods line, which leaves Piotroski's gross-margin signal undefined and is left abstaining; 177 lack only long-term debt) led to two findings. (1) THE HISTORY DEFECT: scripts/history_harvest.kept_tags stored the tags the parser picks values from but not the 43 _BORROWING_CONCEPTS the zero-borrowing rule checks, 28 of which were absent, so in the population test a filer whose debt sat under any of them was read as debt-free while the live scan read it as unknown. kept_tags now carries every borrowing concept; the 28 were back-filled from the cached data sets and 10,804 filers' as-of rows rebuilt. (2) THE LIVE DEFECT: ConvertibleLongTermNotesPayable and UnsecuredLongTermDebt were missing from _BORROWING_CONCEPTS itself, so 8 of the 217 US tiered names read at 0.0 on the 2026-09-28 scan carried a FY2024+ balance there (ServiceNow $1,491M, Zscaler $1,701M, Nutanix $1,349M, iRhythm $650M of convertible notes) and scored as debt-free; they now read unknown on the leverage signal. TESTED AND DECLINED: reading nine noncurrent component concepts as the debt figure (one figure per year, one source per compared pair) would have given 53 more names a pair on live EDGAR, but on the corrected history it lowered the population test (rank-IC 0.243 -> 0.238, deflated Sharpe 0.958 -> 0.907, below the bar); recorded as a trial. CORRECTED CERTIFICATION (history repaired, the two-concept fix applied, held-out halves re-read by scripts/winner_factor_attribution.py): rank-IC 0.244 (was 0.267), tercile spread +0.244 (was +0.300), L/S Sharpe 2.11, t 4.71, Deflated Sharpe 0.966 over N=15 recorded trials (the failed variants are counted), held-out halves 0.250 and 0.219; RANK CERTIFIED. The measured >= +150% base rate on the corrected graded set is 8.4% of 13,714 (BASE_RATE 0.083 kept; within 0.069-0.243). Calibration re-read: still FAILS on one held-out half (Brier skill -0.001, slope -0.54), CALIBRATED stays False. Published five-year quartile figures re-derived: top quartile +55% median, 6.5% lost 70%+ (unchanged); everyone +30% / 18%; bottom quartile -54% / 44% (were -64% / 47%). Emerging overlap under engine 2.0.0: inside the Winner top quartile +71% / 7%, outside +36% / 15%. NVDA's tier history unchanged. Every document quoting these figures updated.
- 1.15.0 (2026-09-28, THE FOUR NAMED WEAKNESSES, MEASURED; ONE DECLINED, ONE DISCLOSED, TWO PENDING DATA): the owner asked what could be done about the rank's weaknesses named after the detector research. (1) THE ANNUAL LAG - DECLINED. scripts/backtest_winner_quarterly.py scores an identical universe at 48 quarter-end dates 2011-2022 in two arms: the latest 10-K row, and the same row with its latest fiscal-year point replaced by trailing-twelve-month flows and the latest balance sheet from the 10-Qs filed since (every 10-Q 2009-2026 harvested with the live parser's tag set, 289,428 filings, each carrying its own year-ago quarter). Three-year rank-IC: annual +0.2197, refreshed +0.2199, mean paired difference +0.0002 (sd 0.011), refreshed ahead in 26 of 48 cohorts; over the four non-overlapping year-end cohorts t = +0.15. NVDA reached Strong candidate two quarters earlier in 2020 under the refreshed arm and otherwise carried the same tier. A quarterly refresh does not improve the rank, so the engine keeps annual inputs (production already splices a vendor TTM point for revenue, income and cash flow). Limits: graded names are those the price feed carries (141-1,954 per cohort, the same in both arms); avg_roic, Altman Z'' and the inflection breakdown kept annual values in the refreshed arm. (2) CYCLICALS AT A COMMODITY PEAK - DISCLOSED, NOT RE-ENGINEERED. On the point-in-time panel the top quartile of Basic Materials and Energy names (n 640) had a five-year median of +29% with 9% losing 70%+ and 4% rising fourfold, against +62%, 6% and 7% for the top quartile of every other sector (n 2,838) - weaker, but still above the population's +26% and far above their own sectors' +3% and 25%. On the 2026-09-27 scan 32 of 161 Strong candidates are in those sectors and none carries the cyclical_peak flag, whose rule (inflection._cyclical_peak) needs two sub-8% troughs and a 5-point give-back: a monotone commodity ramp is not a cycle by its definition, and is a peak only in hindsight. A cyclical top-quartile name historically delivered about half the median gain of a non-cyclical one; the docs now say so and the tier rule is unchanged, because excluding them would drop names that beat the base rate. (3) PIOTROSKI COVERAGE - MEASURED ON THE LAST SCAN, REMEASURE AFTER THE NEXT. Of 1,519 US names carrying a tier on the 2026-09-27 export, 711 have an F-score and 808 do not; the first missing input is gross profit for 434 and long-term debt for 261; that export predates the derived-gross-profit and zero-borrowing rules (b8801135) so the funnel is re-run on the first scan after them before any further fix. (4) NON-US FILERS - LABELLED, CERTIFICATION NOT FEASIBLE WITH THE SHIPPED PARSER. 89 of the 308 TSX names in the universe match a Canadian-incorporated SEC filer by name (63 file 40-F, 11 20-F, 15 10-K); the 10-K filers are already inside the population test, and the 40-F/20-F filers report under IFRS tags the us-gaap parser does not read, so their certification would need a second parser. The panel, guide, glossary and methodology now state that Canadian names are scored by the same rule but were not part of the test.
- 1.14.0 (2026-09-28, A MULTI-BAGGER DETECTOR WAS PRE-REGISTERED, BUILT AND FAILED ITS BAR; FINAL, NOTHING SHIPPED): the owner asked for a mechanism that flags a name at the first sign of gravitating toward a multi-bagger, with reasoning, and that can tell its false positives. The plan was fixed before any result was seen (Drive: Reports/'Multi-bagger Detector - Pre-registered Research Plan 2026-09-28'): a conjunction of ten point-in-time preconditions (eligible by the filer's own public float, small, profitable, cheap, quarterly acceleration, operating leverage, no dilution, insiders net buying, balance sheet, a rising count), a flag = every hard precondition and at least K soft ones with K chosen on three training blocks and read on the two never seen, both ways round; the bar = on BOTH halves >= 2x the base rate of a 4-bagger in five years, wipe-outs no higher than the population's, recall >= 20%, flags <= 2% of the universe; stopping rule = if nothing meets it the conclusion is final. Two free SEC records were harvested for it, tested and point in time: every 10-Q 2009-2026 (289,428 filings, each carrying its own year-ago quarter; scripts/history_quarterly.py) and every Form 3/4/5 2006-2026 (4,458,409 filings, 862,147 open-market purchases; scripts/history_form4.py; SEC data-set page pinned T0, Lakonishok & Lee, Cohen-Malloy-Pomorski and Chan-Jegadeesh-Lakonishok pinned T1). RESULT (scripts/multibagger_detector.py, five configurations recorded in scripts/multibagger_trials.jsonl): none met the bar. Lift on the held-out halves 0-1.9x and never 2x on both; recall never above 1.8%; the best half (1.89x on 50 names) paired with a half where the same rule flagged 19 names of which none multiplied. No single precondition lifts the 4-bagger rate above 1.19x; what they do is remove failures (0 to 5 preconditions met: 4-bagger rate 6.8% to 8.1%, wipe-outs 29% to 8%). Insider buying is two-sided (4-bagger 7.5% vs 6.4%, wipe-outs 33% vs 18%); quarterly acceleration carries no lift (0.95x). NVDA, the owner's example, was a 4-bagger-in-five-years from every filing FY2012-FY2021 from a $5-10B float, never a penny stock; the detector never flagged it (never small), while Winner Odds' own point-in-time percentile placed it in the top 20% from FY2012 and the top 10% in FY2015/17/18/21 - along with ~1,400 other top-decile names whose 4-bagger rate (6.6%) matched the base (7.1%) with a third of the wipe-outs (7.6% vs 19.4%). CONCLUSION: the observable preconditions of a 4x are, at entry, those of an ordinary compounder; a flag from this data would be wrong 88-93% of the time and cannot tell its false positives. Macro was raised by the owner and is outside this data entirely: not tested. The quarterly and insider records remain available as factual context for a thesis, never as odds.
- 1.13.0 (2026-09-28, CAN THE DATA FIND MULTI-BAGGERS? RESEARCH ONLY, NOTHING SHIPPED): a point-in-time panel of 54,965 operating company-years FY2009-2022 (scripts/multibagger_panel.py), every signal from what was knowable at entry, graded at three and five years with bankruptcies and delisted names included, every model judged on blocks it never saw both ways round (scripts/multibagger_model.py; trials in scripts/multibagger_trials.jsonl). Market value is the filer's OWN reported public float (dei:EntityPublicFloat, scripts/history_float.py, 10,063 filers) because shares-as-filed times a split-adjusted price makes future winners look small. FINDINGS. (1) PRICE SIGNALS WERE SURVIVORSHIP: only 5.7% of bankruptcies had a year of pre-entry prices against 85% of survivors; on FY2021-2022, where delisted names carry prices too (Alpaca, production host), 12-month volatility has AUC 0.526 for +150% with 55% of its top decile losing 70%+, and a model trained on survivor prices has lift 1.01x, AUC 0.494, a median pick of -67%. (2) FUNDAMENTALS AND VALUE buy both tails: held-out +300% lift 1.6-2.2x with 32-42% of the top decile wiped out and a negative median; tenfold-in-five-years lift 1.8-2.6x with 41-56% wiped. Consistent with Bali, Cakici & Whitelaw (2011, pinned): lottery-like stocks earn less on average. (3) THE ONE CONSTRUCT THAT HOLDS - small public float, profitable (quality gate), cheap on sales, the recipe Yartseva (2025, pinned) reports on multi-baggers only - beats the base rate on both halves at both horizons with FEWER wipe-outs (3y +150%: 1.44x and 1.54x; worst case with every ungraded name a loser 8.8% vs the population's 5.0%; combined z 2.95) - but its median public float is $11M: it exists in nano-caps only. In investable sizes it fades: $50M-$300M 1.15-1.26x (z <= 1.0), $300M-$2B 1.03-1.28x, above $2B BELOW the base rate. CONCLUSION: this data cannot find multi-baggers among investable companies; the Winner Odds claim stays 'durable compounders and failure avoidance' (its top quartile: five-year median +55%, 6.4% wiped; bottom: -64.5%, 47%). DEFECT FOUND IN THE POPULATION TEST, NOT YET FIXED: the $1 penny guard reads the split-ADJUSTED entry price, so 155 of the 1,391 sub-$1 entries with a known float were real companies with a public float of $50M or more (future split-heavy names); the guard should read the float.
- 1.12.0 (2026-09-28, THE RANK IS CERTIFIED; A PROBABILITY IS NOT): the owner adopted the measured base rate (BASE_RATE 0.083, bounds 0.068-0.245), and with it every gate of the one certification rule holds on the population test - five independent windows, t 4.65, Deflated Sharpe 0.974 over N = 13 recorded trials, rank-IC 0.267, held-out halves 0.277 and 0.234, spread +0.300. That certifies the ORDERING. Whether a probability of winning may be shown was then tested separately and FAILED (scripts/winner_calibration.py, the checks of Van Calster et al. 2019, pinned T1): a logistic of the +150% outcome on the composite percentile, fitted on three blocks and judged on the two it never saw, has Brier skill 0.0000 and -0.0003, AUC 0.505 and 0.492, and calibration slopes 0.48 and -0.79, on both halves; the fitted odds span only 8.3%-9.5%. The reason is in the data, not the fit: the +150% rate is flat across composite deciles (6.8%-9.3%). The composite ranks average outcomes and failures well and says nothing about which names become multi-baggers. CALIBRATED stays False, no odds are shown, the thesis prompt forbids a percentage, and every surface keeps 'rank only'. The honest product claim is 'ranks durable compounders and avoids failures', not 'odds of winning'.
- 1.11.0 (2026-09-28, ACQUIRED AND GONE-DARK NAMES ARE PRICED TO THEIR LAST PRINT; THE BASE RATE IS THE LAST OPEN GATE): the current-ticker feed cannot see a delisted name, so every acquisition and going-dark was a bucket with bounds. MEASURED on the production host (the only place the keys live): Alpaca's IEX feed serves a delisted symbol's bars to its last trading day (Activision 2023-10-13 at 94.44 against the $95 deal; Twitter 2022-10-27 at 53.76 against $54.20) and reaches back only to 2020-07-27; yfinance returns nothing for either; EDGAR full-text search names no ticker for a delisted filer. The ticker comes from the filer's own last 10-K: the inline-XBRL document name (atvi-20221231.htm) or the cover page's 'under the symbol' (scripts/history_tickers.py: 2,363 filers tried, 1,028 resolved for the windows the feed can reach). scripts/history_px_events.py + history_alpaca_job.py (run where the keys live, stdin to stdout, no database, no credential printed): 3,209 requests, 848 of the 1,028 tickers known to the feed, 2,618 requests priced. An acquisition is graded at the last close before the exit (the deal price to within the arbitrage spread, no prose parsed) and a going-dark at its last print, each the terminal value of a window that ends early. RESULT: 838 name-years priced to a delisting exit; graded 13,311 -> 14,149; bounded exits 3,598 -> 3,033 (the rest exited before the feed's first day); rank-IC 0.267, tercile spread +0.300 (+0.438 vs +0.138), tiers monotone in mean under three tiers (Strong +0.435, Watch +0.398, Unlikely +0.218), long/short Sharpe 2.08 over five blocks (t 4.65), Deflated Sharpe 0.974 over N = 13 recorded trials, held-out rank-IC 0.277 (recent half) and 0.234 (early half). THE VERDICT IS NOW 'NOT CERTIFIED: no base rate established' AND NOTHING ELSE. The measured winner rate (>= +150% in three years) is 8.3% of graded name-years, bounded 6.8%-24.5% by the exits that remain unpriced; adopting it as BASE_RATE is the owner's decision, because it is what the probability tiers would anchor on, and fitting the calibration (logistic slope, intercept pinned to that rate, conformal interval) is a further step that turns the product from a rank into odds. Neither is done here.
- 1.10.0 (2026-09-28, THREE TIERS, AND A CERTIFICATION RULE THAT READS HELD-OUT EVIDENCE): TIERS. The population test could not separate the old top decile ('Strong candidate', >= .90) from the 75th-90th band ('Plausible'): by median on both held-out halves they sat level (+0.27 / +0.27 on the recent half, +0.40 / +0.40 on the early half) while the bottom half sat far below (-0.19 / +0.09), and by mean over the whole history Plausible +0.464 beat Strong +0.415. A cut the evidence does not support is a claim the product cannot make, so the .90 cut was REMOVED - not moved, not fitted: three tiers now, Strong candidate = the top quartile (>= .75), Watch (>= .50), Unlikely. Every surface that named the old band (engine cap, screener filter, discovery alerts, badges, glossary, README, guides, methodology ladder) was re-expressed; the boundary buffer and the sticky rule are unchanged. CERTIFICATION. The rule (scripts/_winner_backtest_stats.certify, shared by the basket and population tests) no longer FAILS a model for a rank-IC above 0.15. That ceiling assumed monthly-horizon IC magnitudes; no source was found either way for a three-year horizon (searched 2026-09-28), and survivorship here runs the other way (unpriced names concentrate in the bottom tiers, understating the spread). The owner retired the ceiling, and in its place the rule requires what Bailey, Borwein, Lopez de Prado & Zhu (2015, pinned) argue for: the weighting fitted on three blocks must clear a rank-IC of 0.05 on the two blocks it never saw, BOTH ways round, as measured by winner_factor_attribution.py and written to scripts/backtest_winner_oos.json; a certification with no such file on record is refused, never assumed. The floor of 0.05 on the in-sample IC, t > 3, Deflated Sharpe >= 0.95, five independent windows, a positive spread and a base rate remain.
- 1.9.2 (2026-09-28, ZERO BORROWING IS READ AT THE FILING, AND THE TEST LEARNED TO REFUSE A SUB-PENNY QUOTE): the 245 production names of 1.9.1 tag lease liabilities (182 OperatingLeaseLiability) and facility capacity and 0 of 245 a borrowing balance the parser missed, so the rule moved to the source: sec_edgar emits 0.0 for noncurrent long-term debt in any year the filer tags total Liabilities and NO borrowing concept of any kind (41 concepts: the long-term candidates plus current, short-term, note, facility and convertible lines) carries a value; a borrowing under an unlisted concept keeps the year unknown; only the noncurrent row (the F-score input) takes the zero, the inclusive total behind the Deep Score and the valuation bridge is untouched. ON PRODUCTION (SEC facts fetched for every one of the 349 US operating names with no annual debt point, 346 with facts): 228 now read zero for their latest years, 96 gain a Piotroski F-score, 59 of 1,939 ranked names change tier, 3 rank for the first time. ON THE HISTORY the rule ranked 4,421 more name-years (20,521 -> 24,942) and exposed a data trap: micro-cap shells with sub-penny quotes entered the graded set (Nuo Therapeutics $0.0001 -> $1.50, read as +1,499,900%) and the bottom tercile MEAN went to +3.98 while the rank-IC read 0.263 unmoved. FIX: an entry under $1 is a 'penny' bucket (482 of 24,942), counted, never graded, never an exit - Hou, Xue & Zhang (2020, pinned) for why microcaps must be guarded, the $1 line a disclosed house threshold. FINAL RESULT (zero rule + floor, measured weights, gate live, complete exits): ranked 24,942, graded 13,311 (12,587 priced + 724 bankrupt), rank-IC 0.265, tercile spread +0.282 (+0.445 vs +0.163), long/short Sharpe 1.87 over five blocks (t 4.19), Deflated Sharpe 0.976 over N = 11 recorded trials. Piotroski coverage over voting names 15% -> 32% and its own IC 0.166 -> 0.242 (t 8.8). The shipped weights still hold on blocks never used to fit them: rank-IC 0.290 (recent half) and 0.234 (early half) against 0.223 and 0.172 for equal weights. TIERS: by MEDIAN out of sample the bottom is far below the rest (-0.19 / +0.19 / +0.27 / +0.27 recent half; +0.09 / +0.37 / +0.40 / +0.40 early half) and Strong sits level with Plausible rather than above it; by mean over the whole history Plausible +0.464, Strong +0.415, Watch +0.410, Unlikely +0.237. The top two tiers are not separated by this test and the copy must not imply they are. TRIAL LOG: the five entries recorded while the grading was corrupted by sub-penny quotes were REMOVED (they measured a data error, not a strategy); the log otherwise keeps every scheme ever tried, including refits after each coverage change, which over-counts N and lowers the Deflated Sharpe rather than raising it. STILL NOT CERTIFIED: the rank-IC sits above the house plausibility band and no external base rate exists.
- 1.9.1 (2026-09-28, A DEBT-FREE FILER SCORES ZERO ON PIOTROSKI'S LEVERAGE SIGNAL, NOT UNKNOWN): Piotroski (2000, p. 8, pinned) defines F_DLEVER as one if the ratio of long-term debt to average total assets FELL and zero otherwise; a filer with no borrowing tags no long-term-debt line, so its series was empty and the whole F-score abstained. winner_features._piotroski now reads an empty long-term-debt AND total-debt history - in EVERY point including the vendor's TTM one - with a balance sheet present as zero debt in both years: a 0 on that signal, the other eight still counted. THE TTM POINT WAS THE TRAP: a first cut read only the annual EDGAR points and fired on 256 US names, 241 of which carry a total debt in the vendor's TTM point (AAON $0.45B) - debt the parser's tag list does not read, not debt that does not exist; that cut would have scored levered companies as unlevered. MEASURED ON PRODUCTION (2,891 rows of 2026-09-27, both ways): 7 of 2,052 US operating names gain an F-score (ANET, FNV, GTLB, IRMD, ISRG, POWI, UTHR); 1,939 ranked either way; 5 change tier. Of the 18 US names with no debt point anywhere, SEC facts show 10 with no borrowing concept at all and 5 with a facility or note concept and no balance either source reads. THE 241 ARE THE NEXT STONE: their borrowing concepts are being surveyed so the parser's Long Term Debt candidates can be extended on evidence rather than guessed.
- 1.9.0 (2026-09-28, GROSS PROFIT IS DERIVED WHERE THE FILER NEVER TAGS IT): 810 of 2,583 US operating names on the 2026-09-27 production scan carried no gross profit at all, so Novy-Marx gross profitability and the Piotroski F-score abstained for each. Fetching their SEC facts: 807 have companyfacts, 104 file a GrossProfit tag that did not reach the row (open item), 424 file a cost-of-revenue line, and 393 now get a gross profit as revenue - cost of revenue (sec_edgar.statements_from_facts; a filed GrossProfit always wins; a year with either side unknown stays unknown). The Deep Score already derived its gross margin this way, so its numbers do not move; the Inflection margin trend gains data. ON THE HISTORY, RE-RECONSTRUCTED WITH THE DERIVATION: 7,140 of 60,555 ranked filings gain a gross profit; gross_profit coverage over voting names rises 52% -> 67% and Piotroski 11% -> 15%; ranked+voting name-years 17,962 -> 20,521 (+14%), graded 10,342 -> 11,791. Result with the measured weights and the live gate: rank-IC 0.197, tercile spread +0.159, tiers monotone in mean (Strong +0.524, Plausible +0.465, Watch +0.414, Unlikely +0.388), long/short Sharpe 4.48 over five blocks (t 10.0), Deflated Sharpe 0.983 over N = 6 recorded trials. The shipped weights still hold on blocks never used to fit them (rank-IC 0.214 recent half / 0.168 early half, against 0.164 / 0.124 for equal weights). STILL NOT CERTIFIED: the rank-IC sits above the house plausibility band and no external base rate exists.
- 1.8.0 (2026-09-28, THE WEIGHTS ARE NOW MEASURED, AND TWO FACTORS ARE UNWEIGHTED): until this date the seven core weights were hand-set priors (1.3/1.0/1.0/0.7/0.8/1.1/0.9) that nobody had tested. scripts/winner_factor_attribution.py measured each factor's rank information coefficient against three-year forward returns on the SEC 10-K population, per independent three-year block, FY2009-2022, direction applied: margin_stability +0.250 (t 4.4, measurable on 66% of voting names), roic +0.170 (t 4.5, 40%), piotroski_f +0.166 (t 5.4, 11%), rule_of_40 +0.137 (t 6.1, 72%), gross_profit +0.106 (t 4.9, 52%) - all past the t > 3 bar of Harvey, Liu & Zhu (2016); growth_accel +0.006 (t 0.4, 36%) carries nothing; accruals -0.145 (t -8.9, 71%) runs AGAINST Sloan's direction, consistent with Green, Hand & Soliman (2011) on the anomaly's decay. The weights are the five positive coefficients normalised to sum 1 (Grinold 1989: weight a signal by its IC): margin_stability 0.30, roic 0.20, piotroski_f 0.20, rule_of_40 0.17, gross_profit 0.13; growth_accel and accruals stay core at weight 0 - shown on the scorecard, counted in the five-of-seven presence floor, never in the composite, and NOT flipped (a sign fitted to one sample is not a finding). VALIDATED OUT OF SAMPLE BEFORE IT WAS ADOPTED: fitted on blocks 0-2 and read on 3-4, rank-IC 0.162 -> 0.196 and tercile spread +0.102 -> +0.119 against the priors; fitted on 2-4 and read on 0-1, 0.140 -> 0.173 and +0.137 -> +0.169; top-decile three-year bankruptcy 0.5% vs 3.2% for the rest; Strong's mean return at the top of the tiers on the recent half. Equal weights matched the priors, so the gain is the two dropped factors and the tilt, not the abandonment of judgement. FULL-HISTORY RESULT WITH THE ADOPTED WEIGHTS (in-sample for the weights, gate live, complete exit record): rank-IC 0.191 (was 0.157), tiers MONOTONE in mean for the first time (Strong +0.493, Plausible +0.473, Watch +0.424, Unlikely +0.409), long/short Sharpe 2.85 over five blocks (t 6.38), Deflated Sharpe 0.998 over N = 6 recorded trials (the four schemes tried are in the log); the tercile MEAN spread fell +0.157 -> +0.127 because the old top tercile held more small-cap outliers, and two yearly spreads are negative (FY2017 -0.294, FY2019 -0.168). WHAT DID NOT MOVE: the share of names reaching +150% in three years is flat across tiers under every scheme (top decile 6.6% vs 9.4% for the rest on the recent half) - the ranker finds compounders that survive, not the right tail, and the copy says so. STILL NOT CERTIFIED: the rank-IC now sits further above the house 0.05-0.15 plausibility band, and no external base rate exists. THE BAND'S PREMISE IS ITSELF UNSOURCED for a three-year horizon (it was set with monthly-IC magnitudes in mind); it was NOT loosened to pass, and re-sourcing it is a decision for the owner. ON PRODUCTION (2,891 rows of 2026-09-27, scored offline both ways): 1,939 ranked either way; 616 of 1,939 change tier; Strong 164 -> 171 with 102 names in both; the largest moves are Unlikely -> Watch (176) and Watch <-> Plausible (109 each way). The test itself gained the distress gate for this run: Z'' is now on 43,545 of 79,582 reconstructed rows from two backfilled tags (Liabilities, RetainedEarningsAccumulatedDeficit); it capped 143 name-years out of Strong and 589 into Unlikely and left Strong's mean unchanged.
- 1.7.0 (2026-09-27, the tier question and the missing third — DIAGNOSED, NOT TUNED): Is Plausible > Strong real? Over the five independent three-year blocks Strong beat Plausible in one and trailed in four (differences +0.12, -0.04, -0.18, -0.25, -0.04; mean -0.077, t = -1.20); a bootstrap interval on the pooled gap is [-0.34, +0.02]. MEDIANS ARE MONOTONE (Strong 0.317, Plausible 0.305, Watch 0.281, Unlikely 0.089); the mean gap is a right tail of small-cap Plausible names (small Plausible +1.007, n 391, vs small Strong +0.358, n 348) while large caps order as expected (Strong +0.538 vs Plausible +0.443). Three-year bankruptcy runs 1.5% in Strong to 4.7% in Unlikely (1.5% top decile, 7.3% bottom). The +150% winner rate is FLAT across tiers (8.5 / 9.6 / 8.2 / 9.3%): this ranker separates mean outcomes and failures, not the right tail, and the register says so. DECISION: the cuts stay; fitting them to one uncertified test is what the Deflated Sharpe penalises. THE MISSING THIRD: the full-text exit harvest was incomplete (700 Form 25s in 17 years, none for 2013-2016, against 326 in ONE quarter of the EDGAR full index) and wrong in kind (a Form 25 is not an exit: Ranpak, CIK 1712463, delisted its SPAC units in 2019 and files to this day). Replaced by each filer's own submissions feed (scripts/history_submissions.py; 13,950 of 13,950 filers; T0): an exit is bankruptcy (Item 1.03), ACQUIRED (Form 25 or Item 3.01 with an Item 2.01 or 5.01 within a year) or WENT DARK (Form 15), and NEVER while the filer keeps filing 10-Ks. Re-run: unpriced 6,135 -> 3,852 of 17,962 ranked (34% -> 21%); 1,556 acquired and 978 went dark are carried as bounded buckets (no deal price is parsed), so the graded set is unchanged and so are its statistics (rank-IC 0.157, tercile spread +0.157, Sharpe 1.57, t 3.51); the winner base rate is bounded 7.2%-27.1% by those exits. STILL NOT CERTIFIED for the same two reasons. SURVIVORSHIP RUNS THE OTHER WAY: unpriced names are 24% of Strong and 39% of Unlikely, so their exclusion flatters the BOTTOM and understates the spread. DISCLOSED, NOT FIXED: the reconstructed rows carry no Altman Z (Liabilities and RetainedEarnings were never harvested and the ZIPs are not cached), so the shipped distress gate was INERT in this test - every one of the 1,785 Strong name-years carries solvency_unknown. Fixing it is a re-harvest of two tags.
- 1.6.0 (2026-09-27, after the first production scan on schema v9): A NAME THE GATE COULD NOT JUDGE NO LONGER LOOKS JUDGED. The distress gate flags a Z'' below 1.10 and says nothing otherwise, so an operating company whose Z'' could not be computed - a balance-sheet input missing from its EDGAR frame - ranked with the same clean badge as a name the gate had passed. MEASURED ON THAT SCAN (2,891 rows): Z'' is NULL on 496, of which 382 are the exempt classes (correct: not a Z'') and 114 are gated operating companies with an input missing (21 lack current assets/liabilities, 85 retained earnings, total liabilities or EBIT); two of the 114 ranked Strong candidate. The engine now emits a solvency_unknown flag for exactly that case, rendered as 'solvency not assessed' beside the tier on the panel and the winners table. It is a FLAG, NOT A CAP: a missing input is unknown, not distress, and capping would invent a number. The first scan itself, for the record: 2,195 voting names, 366 abstained for fewer than 5 of 7 core factors (no abstain wave), 550 distress flags against 1,323 under the old 1.8 line, 161 Strong / 239 Plausible / 408 Watch / 1,131 Unlikely, 585 of 1,926 names ranked on both days changed tier. Piotroski is measurable on 797 of 2,052 US operating rows because gross profit is missing on 895 and long-term debt on 536 of them (EDGAR frames without a GrossProfit or LongTermDebt tag); the nine-signal contract was NOT relaxed - the input supply is the lever, and it is not touched here.
- 1.5.0 (2026-09-27, the free survivorship-free history — FIRST POPULATION TEST, NOT CERTIFIED): the ranker was run over every US 10-K filer in the SEC Financial Statement Data Sets (2009q1-2026q2, 2023q1 missing at the SEC's own link), each filer scored AS OF each fiscal year from what it had filed by then, at the EARLIEST filing (point in time), entry at the first bar after the 10-K, three-year horizon, by the SHIPPED engine (scripts/history_*.py, scripts/backtest_winner_population.py). FUNNEL: 64,835 as-of rows -> 17,962 ranked and voting -> 10,346 graded (9,812 priced survivors + 534 bankruptcies counted at -100%) | 155 delisted-other (a bucket with bounds) | 1,326 windows still open | 6,135 UNPRICED. RESULT: rank-IC 0.158, top-vs-bottom tercile forward spread +0.157 (+0.510 vs +0.353), five independent three-year blocks, L/S Sharpe 1.56, t = 3.48, Deflated Sharpe 0.997 over N = 1 recorded trial (nothing to deflate yet). THE PUBLISHED TIERS ARE NOT MONOTONE: Plausible +0.576 (n 1,641) beat Strong candidate +0.435 (n 1,208), which sat beside Watch +0.440 and Unlikely +0.375 - the ordering holds at the tercile level and fails at the top decile. Two yearly spreads were negative (FY2019 -0.159, FY2022 -0.045). NOT CERTIFIED under the one rule (scripts/_winner_backtest_stats.certify): rank-IC 0.158 sits above the 0.05-0.15 plausibility band and no external base rate is established; the gate was NOT loosened to pass. EMPIRICAL BASE RATE, measured for the first time: 8.9%-10.3% of graded names reached +150% in three years (bounds from the delisted-other bucket), and 5.5% if every unpriced name is a loser - a decision for the owner, not adopted here. THE SURVIVORSHIP CLAIM IS PARTIAL AND SAID SO: 6,135 ranked names (34%) have neither a price path nor an exit event, because the SEC's ticker map covers only CURRENT listings and the Form 25 harvest caught 700 delistings over 17 years; those names are excluded and counted, not silently dropped, but a name that quietly delisted is more likely a loser, so every statistic above is an upper bound on the ranker. COVERAGE: only 28% of as-of rows rank, because on EDGAR-only frames Piotroski's nine signals are measurable for ~15% of filers, a gross-profit line for ~50% and ROIC for ~45%; the 5-of-7 floor abstains the rest. The live scan also takes EDGAR frames for US names, so the first nightly scan after Batch 2 is the reach measurement for this floor in production. Parser seams added for the history and inert on the live path (point_in_time, as_of_year, max_years; delta measured 0).
- 1.4.0 (2026-09-27, Batch 4 of the Winner Odds CFA audit): THE EVIDENCE MACHINERY NOW DOES WHAT IT DESCRIBED. scripts/backtest_winner.py described an out-of-fold test and a Deflated Sharpe it did not compute: the leave-one-name-out loop re-appended the same precomputed composites (nothing is fitted, so every metric is in-sample), the 'DSR' was t = SR x sqrt(periods) with no trial count, the entry was priced at Dec 31 of the fiscal year before the 10-K existed, the statements were the latest RESTATED values, and the shipped gates were never applied. Now: statements are taken at their EARLIEST filing (sec_edgar point_in_time=True), the entry is the first bar after the 10-K's filing date (no filing date, no observation), the basket is scored by the shipped engine (score_universe: gates, percentile, tiers) and the PUBLISHED tier is graded beside the composite, the metrics are labelled in-sample, and the Deflated Sharpe is Bailey & Lopez de Prado (2014) - PSR at the expected maximum Sharpe of N trials - with N and V[SR] read from scripts/backtest_winner_trials.jsonl, which records every configuration ever run once; certification requires DSR >= 0.95, t >= 3, 5 windows, the IC band, and an external base rate. THE FORWARD GRADER'S SURVIVORSHIP CLAIM WAS ONE LAYER SHORT: it separated delisted from immature, but the vendor returns NO history for a delisted symbol, so those names landed in 'unpriced' and were excluded; scripts/winner_track_record.py now prices them from the nightly price cache's last bar (closes as stored, not re-adjusted - printed) and counts the terminal loss. BASE_RATE = None: no external rate exists, prob_tier raises rather than assigning a tier from nothing. NOT DONE, stated: the basket is still 38 hand-picked survivors and EDGAR depth still gives ~1 independent 3-year window - certification remains unreachable on free data, and buying history (Sharadar) is a decision for the owner; its price could not be read from the vendor's page (JS-rendered) and is not quoted here. ALSO FOUND WHILE BUILDING THIS, AND FIXED BEFORE THE FIRST SCAN THAT WOULD HAVE HIT IT: US names take EDGAR frames (scanner._deep_statements), and the EDGAR balance frame carried no total assets, current assets or current liabilities - so Batch 2's asset-scaled factors would have been unknown for every EDGAR-sourced US row on the first nightly scan and the 5-of-7 floor would have abstained most of the US universe; sec_edgar now serves Assets, AssetsCurrent, LiabilitiesCurrent and a separate NONCURRENT long-term-debt row (the inclusive concept stays the leverage measure), and compute_metrics reads the noncurrent row first so Piotroski's leverage signal stays long-term debt as published. Each rule hand-checked in tests/test_winner_backtest_stats.py and tests/test_sec_edgar_facts.py; five injection-registry entries watched red.
- 1.3.0 (2026-09-27, Batch 3 of the Winner Odds CFA audit): PEER-RELATIVE, ONE SLOT PER COMPANY, OPERATING COMPANIES ONLY, AND A BOUNDARY BUFFER. Every factor had been z-scored over the WHOLE universe, so the top decile partly measured sector membership: on the 2026-09-25 scan 20.1% of Basic Materials and 14.4% of Technology names were 'Strong candidate' against 1.9% of Industrials and 0.0% of Real Estate and Utilities. Each factor is now standardised within the name's PEER GROUP - its sector for an ordinary company, its business class otherwise, the identical cohort key the Deep Score uses (deep_score_engine._peer_group; a group under 20 voting names uses the universe's moments) - and the composite is ranked ACROSS the universe, so 'Strong' is still the top decile of the field (Seeking Alpha grades every metric within sector and ranks overall the same way, its own FAQ, T2). ONE LISTING PER COMPANY VOTES (deep_score_engine._one_row_per_issuer, the same decider): 108 of the 2,343 ranked rows were a second listing of a company already in the field, and 10 companies held two Strong slots; the other listings now copy the voting row's result and the shortlist collapses them (ranked_batch.leaders_only). NOT RANKED: banks and spread lenders (94), equity REITs (107), regulated utilities (92), insurers (66) and funds (4) - 363 of 2,343 - because Novy-Marx, Piotroski and Sloan are not defined for a deposit-, float- or rate-base-funded balance sheet and none of them could clear the 5-of-7 floor; they abstain with a stated reason and the Deep Score grades each on its own class set. A class-specific Winner Odds set is a DECISION not taken, recorded here. BOUNDARY BUFFER: a name within TIER_BUFFER (0.03) of a cut keeps last scan's ADJACENT tier; a jump across two tiers is never held. MEASURED before this change only: the re-rank is read the morning after the first nightly scan that carries Batch 2's series (none has run yet), so this entry records reach, not the delta.
- 1.2.0 (2026-09-27, Batch 2 of the Winner Odds CFA audit): EACH CORE FACTOR NOW COMPUTES WHAT ITS NAME CLAIMS, AND SAYS SO IN A REGISTRY THE GATE READS. Three factors had shipped as proxies under the convention's own name: 'Piotroski F-score' was seven modified signals on EQUITY denominators with free cash flow standing in for operating cash flow, and a name with a signal unmeasurable scored its unknowns as failures; 'Low accruals (Sloan)' was (NI - FCF)/|NI| - it charged capex as an accrual, exploded near breakeven, and its sign was opposite Sloan's on 585 of 2,343 ranked names; 'Gross profitability' was gross MARGIN, where Novy-Marx (2013) is gross profit over TOTAL ASSETS. 'Margin durability' was current minus average margin - margin EXPANSION, which rewarded a cycle peak. reinvest_runway (ROIC x retention) correlated 0.905 with ROIC, so ROIC carried ~28% of the composite where the table said 17%. FIXED: four per-year balance-sheet series are now stored (total assets, current assets, current liabilities, long-term debt; schema v9), Piotroski is the published nine on beginning-of-year assets and abstains unless all nine are measurable, Sloan is (NI - CFO)/average assets, gross profitability is GP/assets in the same fiscal year, 'Margin stability' is the spread of annual net margin (house, declared), reinvest_runway is removed and the abstain floor is 5 of 7. OVERLAYS: earnings yield is 1/P/E (it read EBIT over BOOK capital - a return, not a yield - and showed Apple at 81%); a PEG or P/E stored as 0 for unknown is unknown (649 of 2,343 ranked names carried PEG 0 and ranked as the cheapest). EVERY factor is declared in winner_model.FACTOR_STANDARD and tests/invariants/test_winner_factor_standards.py runs the Deep Score's Section-12 rule over it. ROLLOUT, stated: on the rows that exist before the first nightly scan the three asset-scaled factors are absent for every name (0 of 2,343), so run_winner_for now REFUSES a cohort where fewer than half the rows carry the series rather than abstain the universe; the re-rank is therefore measured the morning after the first scan, not here. Each formula is hand-checked against its publication in tests/test_winner_features.py; seven injection-registry entries watched red.
- 1.1.0 (2026-09-26, Batch 1 of the Winner Odds CFA audit): THE DISTRESS GATE JUDGED A Z'' VALUE AGAINST THE 1968 Z LINE, ON EVERY CLASS. scanner.py computes Altman's four-ratio Z'' for non-manufacturers (6.56/3.26/6.72/1.05, book equity over total liabilities) WITHOUT its 3.25 constant, and the gate capped any value under 1.8 - the boundary of the ORIGINAL five-ratio manufacturer Z, a different statistic. Altman's own Z'' distress line is 4.35 with the constant, so 1.10 on the stored scale (grey to 2.60; docs/sources/altman-evolution-of-z-score-z-double-prime-zones.txt, T1); the gate had put his whole grey zone inside 'distress'. It also ran on banks, insurers, REITs, utilities and funds, whose balance sheets carry no classified working capital, and a missing balance-sheet line entered the sum as 0.0. MEASURED on the 2026-09-25 production scan, 2,343 ranked names: 994 carried the distress flag and were forced 'Unlikely', 191 of them top-quartile and 72 top-decile on the composite (Cisco, Constellation Software, UnitedHealth, Goldman); 358 of the 474 banks, insurers, REITs and utilities were capped. FIXED: one decider (factor_helpers.ALTMAN_ZPP_DISTRESS = 1.10, shared with portfolio_core's risk profile, which had carried the Z'' zones as literals since before); the gate abstains for spread, insurance, equity_reit, regulated_utility and investment_company (winner_model.ALTMAN_EXEMPT_CLASSES - utilities are a HOUSE exclusion, financials are Altman's), so those classes carry no Safety pillar rather than a pass; and the scanner stores NULL when any input is missing. RE-SCORED on the same rows: distress 994 to 458, 234 tiers move (all upward; 92 from 'Unlikely' to buy-grade), 2,109 unchanged, the composite untouched (0 of 2,343 differ). NAMED, NOT FIXED: 38 top-decile names stay capped, nearly all SaaS with NEGATIVE Z'' (deferred revenue and accumulated deficits: ASAN, SNOW, DBX, DOCU) - software is not an Altman exclusion and needs a decision, not a silent exemption. The rows whose 0.0-built value becomes NULL are not measurable until the nightly scan rewrites them. ALSO IN THIS BATCH: every surface that called the ranking 'validated', 'proven' or 'top-decile' now says it was tested only in-sample on a hand-picked basket (README, methodology, strategy-report methodology, glossary, the thesis prompt, the Telegram digest); the digest printed UPSIDE as '% below fair value'; and the copy that said 'owner+Neha only' now matches the entitlement (Pro and Premium).
- Ships RANK + percentile tier only. It does not publish a probability, by design, while CALIBRATED is False.
- WHAT IS TESTED (superseding the pre-1.5.0 entry that called the ranking NOT validated on a 38-name basket): the RANK's evidence of record is the point-in-time population test (scripts/backtest_winner_population.py, every US 10-K filer FY2009-2022, judged by certify() in scripts/_winner_backtest_stats.py) — certified at 1.12.0 and again on the corrected history at 1.16.0. The 38-name hindsight basket run by scripts/backtest_winner.py (in-sample, a Dec-31 price paired to a 10-K not yet filed) is superseded and is not evidence. The PROBABILITY is not certified: scripts/winner_calibration.py failed both held-out halves on 2026-09-28, so CALIBRATED stays False and no odds are shown. The forward log has not matured.
- The certification path is longer than the documents implied (audit C8).
Entry points & executable coverage
Inflection Engine
inflection · owner vishalsinha · version 3.0.1 · last reviewed 2026-10-07 · inflection.py
Monitoring
WIRED, IMMATURE (2026-09-28) grades matured monthly snapshots against adjusted prices; none have matured
Metric: rank-IC of the live score with the forward return; forward return and wipe-out share per band; Emerging inside vs outside the Winner top tier. Grader: scripts/inflection_track_record.py. Log: winner_predictions (inflection_score + winner_tier columns).
Population test
not certified 2026-10-06 scripts/inflection_program.py · not certified (re-run 2026-10-06) on the corrected history (SIC map, 2023q1 data set): V7 rank-IC 0.163, spread +0.224, t 7.43, halves 0.159 / 0.136 - the ranks still separate - but deflated Sharpe 0.480 over N=13 against 0.95 (scripts/inflection_program_result.json)
Scope: variant V7 (shipped as engine 3.0.0): the rank over US 10-K filers FY2009-2022
Recorded 2026-09-29: certified · V7 rank-IC 0.169, spread +0.248, t 8.12, deflated Sharpe 0.973 over N=13, held-out halves 0.171 / 0.140
Open: the variant set: V5 and V6 are read on four held-out blocks, and V5's Sharpe (8.44) dominates the variance the deflation uses (3.89; 1.44 without V5 and V6) - whether they count on the same footing is under investigation
Validation
population test + invariant The rank: certified 2026-09-28 on the point-in-time population by scripts/inflection_program.py (variant V7, shipped as 3.0.0; weights measured on the same years); NOT CERTIFIED on the 2026-10-06 re-run under the corrected history (deflated Sharpe 0.480 over N=13 against 0.95; the ranks still separate, rank-IC 0.163; 3.0.1, population_test). The engine stays live and disclosed (owner, 2026-10-07). tests/invariants/test_remaining_entry_point_invariants.py — under 3.0.0 the score is the weight-averaged sector-peer percentile over ONLY the weighted factors a name has, re-based to the bands, so a data-poor name is neither penalised nor inflated; fewer than MIN_SCORED known factors abstains as 'Insufficient signal' rather than scoring zero
Assumptions
Limitations
- 3.0.1 (2026-10-07, MODEL UNCHANGED; NOT CERTIFIED ON THE 2026-10-06 RE-RUN - owner decision: disclose and keep running): the certification program (scripts/inflection_program.py, the same 13 variants, every one counted) was re-run on the history repaired for the Deep Score's own test - EXCEPT V1, whose scores are the frozen engine-2.0.0 set the program reads from its 2026-09-28 frame (engine 3.0.0 no longer computes them), regraded on the corrected outcomes only; its Sharpe (0.71) is one of the 13 in the deflation's variance. The SIC map read drug makers as chemicals and aircraft makers as retailers, and the 2023q1 data set had never loaded, so 3,222 FY2022 rows carried a year-late filing date. V7, shipped as engine 3.0.0, is NOT CERTIFIED: rank-IC 0.163 (was 0.169), spread +0.224 (was +0.248), t 7.43 (was 8.12), held-out halves 0.159 / 0.136 (were 0.171 / 0.140) - the ranks still separate - but deflated Sharpe 0.480 against the 0.95 required (was 0.973). Two things moved it and either alone fails the bar: V7's own block Sharpe fell 3.63 -> 3.32, and the variance of Sharpe ratios across the 13 variants rose 2.32 -> 3.89, most of it from V5 (read on four held-out blocks, 5.43 -> 8.44). OPEN: whether V5 and V6, which are read on four blocks, belong in the deflation on the same footing as the five-block variants (without them V[SR] is 1.44). The engine and the Emerging screen stay live; every surface that called it certified now states the re-run result, pinned to scripts/inflection_program_result.json by tests/parity/test_inflection_claims.py.
- Every threshold is a judgement call. They are now DOCUMENTED judgement calls with a stated basis and a date, which is what E-23 asks for, but not one of them has an external authority behind it and no published standard sets any of them.
- POPULATION READ (2026-09-28): the score was recomputed fundamentals-only (F1-F4, F7-F8 where present; no price, insider or analyst inputs) on every as-of recon row of the Winner Odds population test (64,835 rows, FY2009-2022, point in time) and joined to the same five-year outcomes. Year-by-year rank-IC with the 5y return +0.079 (positive in 10 of 11 years; +0.059 at 3y), against +0.140 for Winner Odds on the same rows. Label Emerging (n 441 graded): 5y median +52%, 10% lost 70%+, 5% rose fourfold; No inflection (n 5,923): +37%, 15%, 6%. The edge lives in the overlap with quality: Emerging AND Winner top quartile (n 145) +64% / 7% wiped; Emerging outside it (n 219) +37% / 13%. This supersedes the curated-basket walk-forward (scripts/backtest_inflection.py, survivorship-biased) as the evidence of record; the docs now position Emerging as a lens on the Winner shortlist, weak positive, not a rival rank and not a multi-bagger finder. Live calls are graded forward by scripts/inflection_track_record.py from 2026-09-28 (none matured yet). THE WEIGHT PASS ON F1 AND F6 (2026-09-28, scripts/inflection_weight_pass.py, same population, F6 fed from the Form 4 record 180 days before each filing): F1 growth acceleration (30 points) scores on 37% of rows and its points alone have rank-IC +0.001 (3y) / +0.016 (5y); the 30-point bucket has the WORST outcomes of any bucket (5y median +28%, 18% wiped, against +66% / 11% for 15-29 points), and removing F1 raises the score's 5y rank-IC from +0.079 to +0.134, positive in 11 of 11 years (3y +0.059 to +0.112, 12 of 12). F6 insider net buying (10 points) is ANTI-predictive: points-alone rank-IC -0.169 (3y) / -0.212 (5y), positive in 0 of 12 and 0 of 11 years; names awarded 10 points had a 5y median of -20% with 32% wiped, names awarded 0 (net selling) +48% with 10.5%; raw, insiders net buying -11% / 31% wiped vs net selling +45% / 12%. Adding F6 lowers the score's rank-IC (+0.079 to +0.058). NOT YET CHANGED: a re-weighting is an engine change with its own re-rank delta and register version. NOT CERTIFIED (2026-09-28, later): run through the ONE certification rule Winner Odds passes, on the same population and grading, engine 2.0.0 fails every gate — t 1.68 (need 3.0), deflated Sharpe ~0, held-out halves 0.059 / 0.032 (need 0.05 each). THE PRE-REGISTERED PROGRAM (scripts/inflection_program.py; plan on Drive; 8 variants, every one counted in the deflated Sharpe): peer-relative percentiles of the fundamental trends within year and sector (V2: t 8.09, DSR 0.851); + the standardised unexpected earnings from the 10-Q record in place of growth acceleration (V3); inside the Winner top half (V4: t 2.65 — Emerging adds little among quality names); IC-weighted on training blocks (V5: halves 0.171 / 0.140); the shippable candidates with IC weights on all blocks: V7 (with SUE) rank-IC 0.169, spread +0.248, t 8.12, DSR 0.958, halves 0.171 / 0.140 → CERTIFIED; V8 (no SUE) t 12.62, DSR 0.949 → fails the DSR gate alone. V7 is only moderately related to Winner Odds (Spearman +0.32 on names both score); its rank-IC is 0.170 on the 10,587 rows Winner does not rank, 0.150 / 0.107 / 0.043 in Winner's bottom / middle / top; five-year top quartile +38% with 14% wiped, bottom -31% with 37% wiped, 4x rate flat (7.1% / 6.2%). SHIPPED AS ENGINE 3.0.0 (2026-09-29, owner's go): compute_inflection is now the per-name half (raw statistics + displayed facts, no score); score_inflection_universe is the cross-sectional half, run by scanner.run_inflection_for over the whole latest-date cohort — each weighted factor the name's percentile among its SECTOR peers (a sector-year under PEER_MIN=20 known names ranks against the universe; ties at the average rank), the score the weight-averaged percentile over the factors known (fewer than MIN_SCORED=4 of 6 → 'Insufficient signal'), re-based to the published bands, capped at 50 by heavy dilution or a cyclical peak. Weights = the population rank-ICs scaled to 100 (Dilution 28, Margin expansion 23 — gross and net trends averaged, Operating leverage 14, ROIC 13, Earnings surprise 12, Cash-flow inflection 10); Growth acceleration, Price breakout, Smart money, Forward growth and Estimate revisions stay in the breakdown as facts with max 0. The earnings surprise is read at scan time from the filer's own companyfacts quarters (sec_edgar.quarterly_net_income: each 10-Q's ~91-day quarter, latest filing wins; Q4 derived as the 10-K year minus the three filed quarters only when all three exist) and stored in fundamentals.earnings_surprise (schema v10). RE-RANK DELTA on the 2026-09-28 book (2,891 rows; SUE known for 1,767): scored in both engines 2,319, label changes 969, Emerging-or-better 109 -> 123 (69 leave, 83 enter); 3.0.0 label mix Inflecting 21, Emerging 102, Building 393, Early/mixed 766, No inflection 1,079, Insufficient signal 220. _REBASE_KNOTS re-derived on that book: raw 50.0/65.0/79.0/91.0 -> 35/50/65/80. The certification is V7's on the fundamentals-only score; the live grader (scripts/inflection_track_record.py) accrues against it from the next monthly snapshot. APPLIED 2026-09-28 (version 2.0.0, owner's go): growth acceleration max 30 -> 10 with the four rungs kept as fractions of the registry max; insider buying max 10 -> 0, still in the breakdown as a displayed fact (scored when known, never a point); the score re-based to the book's band shares (_REBASE_KNOTS). On the population the exact configuration lifts the 5y rank-IC from +0.058 to +0.096 (10 of 11 years positive) and the 3y from +0.045 to +0.078 (11 of 12). Re-rank delta on the 2026-09-27 book (2,423 scored rows): 559 label changes, the Emerging-or-better set 109 -> 111 names (39 leave, 41 enter); no factor definition changed, only its weight.
- 3.0.0 G-SCORE PASS (2026-09-29, registered before the run, DECLINED — V7 STAYS): Mohanram's eight growth-firm stability signals (GSCORE; the Hou, Xue & Zhang 2020 definitions) were read as continuous sector-year percentiles and added to V7 through the same program, N=11 trials: V9 = V7 + the six the record held (ROA, cash-flow ROA, accruals, earnings variability over 16 quarters of 10-Qs, sales-growth variability, capex intensity), V10 = V9 + R&D and advertising intensity (back-filled into the 10-K record that day; R&D known for 32.9% of rows, advertising 3.2%), V11 = the eight alone. The bar, fixed in the script's docstring beforehand: certified AND both held-out halves at least V7's. V9: rank-IC +0.331, spread +0.403, halves +0.379 / +0.286, t 4.62 — but deflated Sharpe 0.169 against 0.95 required, because its year-block spreads swing from +0.17 to +0.60 (V7: +0.18 to +0.34); V10 identical (R&D, advertising and accruals earned weight 0 — their rank-IC is not positive); V11 the same shape (DSR 0.083). WHY THE GAIN IS NOT WANTED EVEN WHERE IT IS REAL: V9 correlates +0.69 with the Winner percentile (V7 +0.32) and reads +0.069 inside the Winner top quartile — the profitability LEVELS (ROA +0.296, cash-flow ROA +0.289) and earnings stability (+0.288) are Winner Odds' quality signal in different clothes, and Emerging Compounders exists to read CHANGE that Winners does not. The eight signals alone find failures, not the tail: bottom quartile 3y median -55% with 43 in 100 wiped, top +32% with 3 in 100 — and the top quartile rose fourfold LESS often (1.3% vs 3.5%). Nothing shipped; the record now carries R&D and advertising; V7 at N=11 reads DSR 0.966, still certified. Run: scripts/inflection_program.py (V9-V11), report on Drive.
- 3.0.0 INSIDER PASS (2026-09-29, registered before the run, DECLINED — INSIDER BUYING STAYS A DISPLAYED FACT): the one published reason raw insider buying might mislead is that most insider trades are routine. Cohen, Malloy & Pomorski's classifier (NBER w16454: classified only with a trade in each of the three preceding years; routine if one calendar month was traded in all three; opportunistic otherwise) was applied per insider and issuer to the Form 4 record, which now carries the reporting owner's CIK (scripts/history_form4.py, 3,502,773 trades, 140,420 insiders). Officer and director trades: 323,294 routine, 256,967 opportunistic, 2,366,669 unclassified (no three-year history). Signal: (buys - sales) / (buys + sales) in dollars over the Form 4s filed in the 365 days before the 10-K. RAW: rank-IC -0.106 (halves -0.102 / -0.123; net buyers 3y median -1.2% with 22.7% wiped, net sellers +24.9% with 8.8%). OPPORTUNISTIC ONLY: rank-IC -0.030 (halves -0.060 / -0.077; net buyers +6.9% with 17.3% wiped, net sellers +32.3% with 4.4%). The filter removes most of the damage and none of the sign: a non-positive rank-IC earns weight 0, so V12 and V13 are V7. HORIZON CAVEAT: the paper measures the month after the trade; this test measures three years after the 10-K, so it does not contradict the paper — it says the signal is not a three-year one. V7 at N=13 reads DSR 0.973.
Entry points & executable coverage
ETF quality score
etf_score · owner vishalsinha · version 1.1.0 · last reviewed 2026-09-16 · etf_scoring.py
Monitoring
NOT MONITORED the mix was measured once on 2026-09-16
Metric: label mix per asset class after each ETF scan. Grader: none. Log: etf_fundamentals.
Population test
not yet tested 2026-10-06 · no point-in-time ETF history exists in this system (no archived categories, expense ratios or holdings by date), so no population test can be built today; the label mix was measured once (2026-09-16) and nothing grades the labels forward
Validation
example tests/test_etf_scoring.py — asset-class corpus over the real universe's hard cases, per-class ladder golden, unknown-window re-basing, scanner keeps a missing window None; each seen RED against its injection
Assumptions
Limitations
- Asset class is inferred by keyword, so a Canadian fund whose name says neither bond, cash, bullion nor portfolio is equity by default; the corpus test names every hard case in the current universe (SGOV is cash before bond, XGD is gold miners not bullion, XEQT is equity not balanced).
- No benchmark-relative return: a bond fund is judged against a fixed ladder, not its index. Bond and cash ladders are judgement calls sized to the 2015-2025 rate regime.
- Commodity and alternative funds are judged on the equity ladder; no lower bar is defensible for them.
Entry points & executable coverage
Position sizing
sizing · owner vishalsinha · version 1.1.1 · last reviewed 2026-10-06 · sizing.py
Monitoring
NOT MONITORED
Metric: realised concentration vs suggested band. Grader: none. Log: none.
Validation
invariant tests/invariants/test_model_entry_point_invariants.py — band/tier laws, one-tier valuation move, AVOID is never promoted, unknown quality is not sized as average
Assumptions
Limitations
- 1.1.1 (2026-10-06, REGISTER ONLY): the Deep Score's population test certified the rank of its 85-point fundamentals core (deep_score 2.1.3). It did not test these bands, the label cut-offs that set them, or the Valuation pillar, so the bands remain untested and every surface that shows one says so.
- 1.1.0 (2026-09-30, REGISTER CORRECTED; the code moved earlier): the client's declared single-name ceiling now REACHES suggest_size (max_position_pct, passed by api/stocks.py and api/risk_api.py) and narrows the band, never widens it. The earlier entry saying the ceiling was never passed to it was stale and is withdrawn. This model's version was never recorded before: '1.0.0' was the builder's default, which it now refuses. The bands are built on the Deep Score labels, which have not been tested against later returns (deep_score 2.1.2), and every surface that shows a band says so.
- Sizes by score tier; ignores volatility, correlation and liquidity (audit D1).
- Bands top out at 12%, against a private-bank single-name cap of 5-10%.
- Tiers share LABEL_CLIFFS with the Deep Score verdict, so a change to the score's scale moves suggested position sizes.
Entry points & executable coverage
Decumulation plan
decumulation · owner vishalsinha · version 1.5.0 · last reviewed 2026-10-06 · decumulation.py
Monitoring
NOT MONITORED
Metric: plan versus actual withdrawal and balance. Grader: none. Log: none.
Validation
invariant + independent oracle tests/invariants/test_retirement_horizon.py holds the plan to the pot AT RETIREMENT (hand-checked: C$1M at 3% real for 20 years is C$1,806,111), proves an already-retired plan moves by nothing, and proves a past retirement age cannot DIVIDE the pot; tests/invariants/test_model_entry_point_invariants.py replays the withdrawal year by year — it must neither deplete early nor survive an extra year; test_remaining_entry_point_invariants.py pins the CPP/OAS break-even against an independent simulation and the withdrawal SEQUENCE against the order the docstring promises, including that being near the OAS clawback annotates the TFSA step without promoting it
Assumptions
Limitations
- 1.5.0 (2026-10-06): THE PLAN LASTS TO THE CLIENT'S FP CANADA SURVIVAL HORIZON AND RUNS NET OF FEES. (a) HORIZON: planning_horizon reads the 2026 PAG section 4(e) Probability of Survival table (CPM2014 + CPM-B, stored in tax_rules.json global.pag and checked row by row against the pinned source) and plans to the age where the chance of survival is 25%: the M or F column, the couple column when a partner's birth year is on file (or a linked spouse's), the most conservative column consistent with what is known when sex is not on file, the later of two 5-year rows between them, and a younger partner extending the plan by the age gap. The house 95 it replaces was below the 25% age for every woman and every couple in the table. A +-5-year sustainable-spend sensitivity is published beside it. (b) FEES: the typed real return is GROSS; retirement_sim.return_basis deducts the pot's book-weighted fund MER (advisor_engine.book_fund_costs) plus any declared advisory fee, and the plan runs on the net figure; with a fund's MER unknown it runs gross and says why. LIMITS: the fee is today's book carried forward; the TER is not included; the couple plan is still a single-person cash flow (no survivor benefits, no pension splitting).
- 1.4.0 (2026-09-30): A PLAN WITH NO YEAR TO FUND IS REFUSED. A retirement age at or after the planning horizon (95) left no year between retirement and the horizon, so the 'sustainable spend' was the whole pot in one year. build_plan now raises 'no_horizon' instead of showing a figure (audit 2026-09-30). The sustainable spend this engine publishes is the MIDDLE case: it lands the pot at exactly zero at 95 only if returns average the assumed real rate every year; when returns vary, spending at that level can run out early, and under the plan's own simulation it lasts in roughly half of paths. It is not a safe spend, and the retirement panel now shows its odds beside it (Retirement Monte Carlo 1.4.0).
- 1.3.1 (2026-09-29): CORRECTION to 1.3.0. It said the owner 'was shown C$13,697 short'. The figure came from a hand-built book that missed his fund holdings - on the engine's own holdings total the old formula gives about C$12,395 - and what the panel displayed was never read, only inferred. The defect and the fix are as described; the number and the verb were not.
- 1.3.0 (2026-09-29): A BENEFIT THE CLIENT HAS NOT ENTERED IS UNKNOWN, NOT ZERO. build_plan computed total income as spend + (cpp_annual or 0) + (oas_annual or 0) and took the gap from it, so a plan with no estimates on file told its owner how far short they were on the assumption that they would receive no CPP and no OAS for the rest of their life. Found by using the product as an adviser would: the owner set a C$100,000 spending target with no estimates on file, and on his engine-measured holdings the old formula states a shortfall of about C$12,395. Measured on production before the fix: 7 users -> 2 with a plan -> 1 with a spending target -> 1 with both estimates missing. Total income and the gap are now null while EITHER estimate is missing, benefits_missing names which, and income_on_file_cad carries the sum of what is known. For a plan with both estimates on file nothing moves.
- tests/test_decumulation.py CERTIFIED the defect: it asserted that total income EQUALS the sustainable spend when no benefit is on file, which is only true if the client receives nothing. It was green throughout. Rewritten to assert the absence.
- STILL LATENT, measured at zero reach and named rather than built: with OAS on file and CPP missing, the clawback estimate and the forced-RRIF test still add CPP as zero, which understates both. No user has either estimate on file (0 of 7), so there is no instance to move; benefits_missing is on the plan for any surface that renders them. The Monte Carlo likewise runs on the portfolio alone while a benefit is missing, and its assumptions block now says so.
- 1.2.0 (2026-09-25): THE PLAN IGNORED WHAT THE CLIENT IS STILL SAVING. `profile_annual_contribution` has been on the KYC form and read by eight surfaces for months; this engine was not one of them, so the projection covered today's pot and silently dropped every future dollar. MEASURED ON PRODUCTION: 1 of the 2 users with a retirement plan has one on file, at C$60,000/yr, which over their 20 remaining years compounds to C$1,612,222 - MORE than their existing pot grows to. Ignoring it was not a conservative rounding, it was most of the answer. Contributions now compound as an ORDINARY annuity (level, paid at each year END: FV = C x ((1+r)^n - 1)/r, C x n when r is 0), matching the convention retirement_sim.goal_success_probability already documents so the two engines cannot contradict each other - the withdrawal side uses annuity-DUE for the opposite and equally deliberate reason, and annuity-due here would overstate by a factor of (1+r). They land pro rata across the sleeves already in use, which leaves the OAS-recovery apportionment ratios exactly where they were; with no book at all they go to non-registered, the sleeve that feeds the clawback test, so the assumption errs toward warning. RETRIEVED 2026-09-25 (T0): the 2026 FP Canada / Institute Projection Assumption Guidelines mention contribution timing ZERO times in 56,153 characters - they govern RATES, not projection mechanics - so no external convention applies to this choice and the repo's own is matched instead. The same document states three times that management fees must be subtracted to obtain the net return; our rate is gross, so the contribution now compounds that overstatement on a larger base, disclosed on the panel and in the methodology.
- STILL NOT MODELLED: any CHANGE to the contribution rate - it is assumed level and real to retirement.
- 1.1.0 (2026-09-25): THE PLAN WAS COMPUTED AS THOUGH RETIREMENT BEGAN TODAY. Every figure derived from the pot used the balance as it stands now, over the drawdown window only, so the years between now and the retirement age were silently dropped. Correct for someone already retired; wrong for everyone else. MEASURED ON PRODUCTION 2026-09-25: 7 users -> 2 with a birth year (the plan renders at all) -> 2 whose retirement is in the future, i.e. the entire population that sees a plan. A 43-year-old retiring at 65 had her sustainable spend understated by 92% and a 45-year-old by 81% (C$49,533/yr per C$1M held against C$89,462/yr), and the SIGN of the headline answer moved: the hand-worked case went from 'C$5,334/yr short of your target' to 'C$26,610/yr ahead of it'. The pot now compounds at the plan's own REAL return over the accumulation years, and the figures that are statements about the retirement years move with it - the assumed non-registered income behind the OAS recovery tax, and the RRIF minimum at 71. What does NOT move is what a client can tie to a broker statement: investable_cad and the sleeve split stay observations of today, published beside investable_at_retirement_cad so the two reconcile on screen.
- Flat real annuity with a fixed withdrawal order, not a year-by-year tax-aware plan (audit E1).
- No pension splitting, GIS interaction, or s.146(8.8) estate treatment.
- CPP/OAS break-even ages ARE computed — the audit's claim that they were missing was refuted.
Entry points & executable coverage
Strategy Advisor — signal rule engine
signal_rules · owner vishalsinha · version 1.3.0 · last reviewed 2026-10-01 · signal_engine.py
Monitoring
NOT MONITORED
Metric: forward outcome of TRIM/SELL signals. Grader: none. Log: none.
Validation
invariant tests/invariants/test_advisor_signal_invariants.py + test_advisor_absence_invariants.py — soft-reason tax guard, verdict never silence, tax always stated, one decider, unknown is never a number
Assumptions
Limitations
- 1.3.0 (2026-10-01): A HOLDING WITH NO PRICE IS NOT SIZED. value_cad joins GUARDED_INPUTS and every rule that reads a weight or a value declares it, so a holding the price feed could not value (no live price and no stored close) gets the not-assessed verdict instead of a size judgement. The ADD rules read its unknown weight as 0% and told the client it was 'only 0.0% of portfolio. High conviction warrants higher weight' (independent review, 2026-10-01). Enforced by tests/invariants/test_signal_aggregate_weight.py's rule-input sweep.
- 1.2.0 (2026-09-30): A POSITION IS SIZED AS A WHOLE. Every size judgement now uses the whole position - the same ticker summed across all accounts - not one account's slice; a trim is still sized in the account it happens in, and its text names that account. No ADD fires on a name whose whole position is at or above its ADD level: the client's declared single-name limit x 40/45 (the house limit 25% -> about 22.2%), the same headroom the house sector pair uses (ADDs stop at 40%, trims start above 45%). The mid-quality trim takes the whole position to 5% by taking the same share off every slice. Today's concentration card fires at the client's declared single-name limit, else the 25% house limit; it was a hard-coded 25%. Buy, sell, trim and add cards on Today now come only from this engine: the AI Tax & Structure Audit contributes structural moves only (MIGRATE, CONSOLIDATE, HARVEST), and every card names its source (audit 2026-09-30).
- Rules are point-in-time on a single holding plus portfolio aggregates; no rule sees correlation, liquidity or realised volatility.
- Concentration limits sit above private-bank policy norms by design (review B5); a client who wants a bank's limit must declare one (max_single_position_pct, max_sector_pct).
- A declared limit LOOSER than the house one is honoured, not refused: the report states the departure for every sector and name concerned but does not re-ask. It records no acknowledgement from the client beyond the declared number itself. The statement is in the report and on the web Strategy page; the signals list, the Android app and iOS do not carry it.
- The sector ceiling is only as good as the look-through beneath it: a fund the ETF scan has not described is unclassified and outside every sector's weight.
- No rule's forward outcome is recorded, so no TRIM or SELL has an out-of-sample hit rate.
Entry points & executable coverage
Strategy Advisor — orchestration, suitability, tax allocation, AI validation
strategy_advisor · owner vishalsinha · version 1.16.0 · last reviewed 2026-10-06 · advisor_engine.py
Monitoring
LOGGING picks recorded since 2026-09-10, none matured; not yet graded
Metric: forward return of diversifier picks vs SPY, by cohort. Grader: scripts/advisor_pick_track_record.py. Log: advisor_picks.
Population test
not yet tested 2026-10-06 · the diversifier picks are logged (advisor_picks, since 2026-09-10) and graded forward by scripts/advisor_pick_track_record.py; no cohort has matured, and the pick logic depends on each client's own book, so no retrospective population test exists
Validation
invariant tests/invariants/test_advisor_*.py, test_fund_lookthrough_invariants.py, test_stress_replay_invariants.py, test_risk_questionnaire_reaches_the_advisor.py, test_growth_and_concentration_disclosure.py — absence never a number, registered-first allocation, AI figures reconciled to the engine, fund fold-through on both spellings of absence, and the two engine-enforced disclosures: the threshold is the pillar's own weight share, an unmeasured share is silence and never 0%, the candidate whitelist carries the measurement through to the report, and no disclosure makes a forward claim the score has not been validated for (CFA V.A), each proven red by 16 entries in tests/parity/test_gate_injections.py
Assumptions
Limitations
- 1.6.0 (2026-09-24): ADVISORY SPLIT FROM RESEARCH - A SINGLE NAME IS NO LONGER SIZED AS ADVICE. Measured on production 2026-09-24: 66 stored reports -> 48 carrying diversifier recommendations -> 180 single-name recommendations -> 180 of 180 carried a suggested_weight. Not a sample: every single name this product has ever shown a client arrived with a position size, behind a rank with no graded forward return. A size is the sentence that turns research into advice - 'consider DLO' and 'put 3-5% into DLO' are different acts, and the second implies a basis that does not exist. The size is now stripped at the engine choke point and the recommendation states how much graded evidence stands behind the screen, INCLUDING none. AN ETF KEEPS ITS SIZE: allocation rests on a declared policy, a tax rule or an arithmetic identity, and this is a split rather than a retreat from advice the product can defend. ENFORCED BY THE ENGINE, not the prompt - the prompt already asked for this shape and produced 180 of 180 the other way. GUARDED IN BOTH RENDERERS TOO, because the 48 stored reports still carry a weight and are re-read long after they were written. No contract change: suggested_weight is absent from contract.py and neither mobile client has a diversifier surface.
- 1.7.0 (2026-09-24): "DIVERSIFIES" NOW SAYS WHAT IT RESTS ON (B2a), AND IS MEASURED WHERE THE HISTORY ALLOWS (B2b). The prompt instructed the model to write '<sector> - diversifies the overweight' while the engine handed it NO correlation input at all: a candidate carries ticker, name, sector, deep_score, the growth-pillar fields and valuation fields, and nothing about how the name moves with the book. The only basis the screen had was that the candidate sits OUTSIDE the overweight sectors. That is a real basis and a weak one - two different GICS sectors can move together closely, which is exactly why the Risk X-Ray measures effective BETS rather than counting names. Every single-name recommendation now states the basis, and where price history allows it carries the measured change in effective bets, computed through the SAME risk_data/risk_concentration pair api/risk_api.py uses so the report and the X-Ray cannot disagree about one book. A name that measurably does NOT diversify is disclosed as such; the recommendation still stands, and the contradiction is stated beside it.
- B2 COVERAGE IS A MINORITY AND THE DISCLOSURE NAMES ITS OWN: measured on production 2026-09-24, 10 distinct live diversifier candidates -> 4 carry the >=60 days of price history a correlation needs. The plan specified this 'on every candidate'; that premise failed on the data. The other 6 say 'NOT measured' explicitly rather than staying silent, because silence beside a diversification claim reads as assent.
- 1.7.0 also adds B6's SUSPENSION RULE, and the honest fact about it is that it CANNOT FIRE YET. advisor_pick_track_record.py grades logged picks against SPY and the name's own sector, but it is network-bound and not in CI, so no graded cohort verdict is stored anywhere the report can read. The rule therefore returns ACTIVE with an explicit basis ('no graded cohort yet' / 'cohorts have matured but none has been graded') rather than ACTIVE with an implied all-clear - a STATE, not a boolean, because 'not suspended' and 'not yet judgeable' are different facts and a boolean collapses them into the reassuring one. When it does fire it REMOVES single names from the report rather than relabelling them; ETF recommendations survive, because allocation rests on a declared policy, a tax rule and an arithmetic identity and is not what failed.
- _SLEEVE_SUSPEND_AFTER_LOSING_COHORTS = 3 is a DECISION, not a fitted number. With no matured cohort there is nothing to fit it to, and a threshold fitted later to the very data it judges would be circular. Three is slow enough that one bad quarter does not silence a screen and fast enough that a persistent loser does not run for years. REVISIT when the first cohorts mature (first grading 2027-01-15).
- WHAT THIS DOES NOT DO: it does not make the single-name screen validated. The pick log and the grader exist; no cohort has matured, and the report now says so with the date the first one will. Nothing here converts a rank into evidence.
- TERMS AND PRIVACY WERE DELIBERATELY NOT TOUCHED. The plan had this batch editing docs/terms.html and docs/privacy.html, which state what the product represents itself to BE. F2 (registerable advice) is parked pending securities counsel, so a drafted legal representation would be the one claim in this programme with no gate behind it and the only one with legal rather than financial consequences. OWED: counsel review, then the wording.
- Recommendations carry no out-of-sample evidence: no pick, TRIM or SELL has a graded forward return yet (CFA Standard V.A). Nothing may describe the output as validated until scripts/advisor_pick_track_record.py grades a matured cohort.
- Suitability is an analytics analogue, not a regulatory determination (CFA III.C): below the six-field minimum every action is review-only with no timing and the missing fields are named; on a complete profile the four declared constraints narrow the plan and never widen it. Measured 2026-09-22: 1 of 6 client books had the four constraint fields on file.
- The declared allocation policy is enforced (a breach leads the plan) for clients with an IPS on file — 2 of 7 users on 2026-09-22. For the rest the questionnaire's mix is shown as a PROPOSAL awaiting acceptance and drift is never measured against it.
- Goal feasibility uses FP Canada's 2026 Projection Assumption Guidelines NET of the book's fund MER and any declared advisory fee when every MER is on file (gross, and labelled with the reason, when not; Batch 5, 2026-10-06) and an i.i.d. normal draw; the basis mix (declared / proposed / currently held) is named on every figure, and the currently-held basis is the return of what the client owns, not of a suitable allocation.
- Sector look-through covers only funds the ETF scan has described (sector_weights on file); the remainder is reported as unclassified, never as a sector.
- The written narrative is an LLM under constraint (schema, ticker and share checks, figure reconciliation, temperature 0); its prose is not deterministic across model versions.
- 1.16.0 (2026-10-06): A LOW SCORE IS A REVIEW, AN UNKNOWN BETA IS NOT A NUMBER, A CDR IS ITS UNDERLYING (audit HIGH residuals P004, P001, P002). (a) PROMPT: the narrative prompt said 'SELL is only for genuine deterioration: AVOID-rated quality (score < 40)' and 'Any holding with score < 40: recommend SELL or TRIM', a sale on a percentile the engine refuses (avoid_fundamental_rank_only). Both now defer to the engine; every holding in the prompt carries engine_actions (the signal list is capped at 40 and a rank-only HOLD is the first thing it cuts); the model is told the Deep Score and its labels have NOT been tested against forward returns. _validate_narrative corrects a model SELL/TRIM of a name the engine held on its rank alone to HOLD, recorded in _validation. Gate: tests/invariants/test_ai_prompts_defer_sales.py reads the RENDERED prompt. (b) STRESS ROWS: the broad-bear and severe-recession rows applied the covered-average book beta to the WHOLE book, so a holding with no beta was shocked at the average. They now go through the page table's decider (_stress_population / _beta_shock_loss): a holding with no beta is excluded and named (stress_unmodelled, stress_unmodelled_note), the loss and loss % are over the modelled holdings, after_cad is not stated while any holding is unmodelled, and with no beta anywhere the row states no loss and no value after ('no betas on file — not stated'; a flat index drop until 2026-10-07). Betas are clamped to [0, 3] here too. The suitability drawdown conflict names the population its loss % is over. (c) LOADER: a CIBC CAD-hedged CDR on a US company carries its underlying's beta (cdr.BETA_FROM_UNDERLYING, 31 receipts checked against CIBC's CDR Directory, T2, retrieved 2026-10-06); on production every no-beta holding was such a receipt (2 of 6 users). Other receipts in the price map are not covered and stay without a beta. SIDE EFFECTS, measured on a test book through the real engine (tests/invariants/test_cdr_beta_side_effects.py): late_cycle_trim and deteriorating_outlook_trim now fire on a covered CDR in a registered account (they stood down for lack of a beta) and still stand down on a taxable winner; recession_risk_highbeta does not fire (it also needs a Deep Score, which a CDR lacks); an ADD of a covered CDR is now judged against a declared drawdown limit instead of listed as unjudged. (d) GOAL RISK: the Strategy report's callout no longer picks a worst case over market-wide rows whose value-after is not stated; it says the shortfall cannot be stated and names the holdings left out. (e) ONE RULE FOR AI TRADE VERBS: signal_engine.ai_action_shown — a model SELL, TRIM or ROTATE stands only when the engine raised a reduce for the ticker, and never as a stronger verb (an exit where the engine raised a trim is shown as the engine's TRIM; SELL vs ROTATE shows the engine's; the model's verb is kept as provenance); matched per (ticker, account) slice, since the engine raises a trim on the slice it allocates it to — a sale on another slice is a review naming where the engine acted; a model ADD/BUY of a name the engine is reducing is HOLD (buys are otherwise ungated: fund and idle-cash advice needs them); an all-cash book is cash_only on the report as on the page (one decider, _stress_basis); a ValueError inside the Tax audit is a 500 unless it is the builder's TaxAuditNoBook (the first version keyed on the rank-only HOLD's rule name, which summarise_signals can hide behind a stronger HOLD); the Tax audit prompt prints the signal engine's verbs, not auto_recommendation's, and labels an investor override as theirs.
- 1.15.0 (2026-10-06): WHAT THE FUNDS COST IS READ ONCE, IN PERCENT, AND EVERY PROJECTION DEDUCTS IT. book_fund_costs is the one MER reader: etf_fundamentals.expense_ratio is a FRACTION and the cost disclosure had published it as weighted_mer_pct unscaled (a 0.20% fund reached the prompt as 0.002) and weighted over fund value only; it is now a percent over the whole book with shares and cash at 0%, with a C$ estimate per holding and per account (MER only; the TER the dealer's NI 31-103 s.14.17 report adds is not included, and the report says so). Reach on production 2026-10-01: 2 of 6 users with holdings hold a fund with its MER on file. _policy_expected_return deducts the book fee (plus any declared advisory fee) via return_basis; the retirement lens carries the 90%-success spend, the fee basis and the FP Canada horizon.
- 1.14.0 (2026-09-30): FOUR FIGURES THE REPORT QUOTES WERE CORRECTED AT THEIR SOURCE (audit 2026-09-30). (a) STRESS TABLE (beta_shock_table): cash is carried unshocked and the loss is over the invested holdings only; each holding falls by its beta x the shock, beta clamped to [0, 3], each holding floored at zero; a holding with no beta takes the covered holdings' value-weighted average beta, never an assumed 1.0, and a book with no betas at all is labelled unadjusted. (b) CRISIS REPLAY AND CVaR (_historical_stress): US-dollar holdings are converted to Canadian dollars at the same day's USD->CAD close before any return is taken; a US holding with no FX history is excluded and named, never priced as if it were Canadian. (c) PERFORMANCE: under one year of history nothing is annualised (GIPS 2020, 2.A.12) and the prompt is given cumulative figures for the period instead; a benchmark whose first close is more than 7 days after the first transaction is not compared; the dollar gap treats the client's dividends as reinvested in the index from the day received, and it measures a gap, not skill. (d) MARGINAL RATE: the rate a taxable trim is estimated at is derived on every read from the client's province and income band, computed from CRA's 2026 brackets at the band midpoint, instead of a snapshot saved with the profile; a client near a bracket edge can differ from it.
- 1.13.0 (2026-09-29): A BREACH ALREADY SHOWN IS A QUESTION ABOUT THE POLICY. A declared policy breach led every report with the same REBALANCE, report after report, and nothing asked whether the policy still described what the client wanted. CFA Standard III(C), 'Updating an Investment Policy' (T0): an IPS is reviewed at least annually, and a client who holds an off-policy book is discussed with - the trade is made or the policy updated. _policy_review asks when (a) the book was outside THIS policy in the same class and direction in the report before, and every report back to breach_since, with the policy not saved since; or (b) the policy was last saved more than _POLICY_REVIEW_DAYS ago. Only reports stored after the policy was last saved count, so re-saving it starts the clock again; a report inside the band or with a failed lens ends the chain; an unknown save date triggers nothing; a failed history read is reported, not treated as no history. The question is appended to the breach action and printed in the policy panel of the report and the web Strategy page; it never suggests moving the policy to match the book. Also: a REBALANCE the model wrote itself is now REPLACED by the engine's, which carries the figures and every engine sentence; keeping the model's used to suppress them (0 of 4 stored breach reports were affected). Reach on production 2026-09-29: 7 users, 2 with a declared policy, 2 whose last two reports were both in breach; the owner re-saved his policy after his last report, so only one client's next report would ask.
- LIMITS OF THE POLICY REVIEW. (a) Reports are generated on demand, so 'already shown' means one earlier stored report, not a period of time. (b) Only the report asks; the Today card still shows the rebalance every day without the question, because its path must not read report history. (c) It does not record the client's answer: re-saving the policy is the only acknowledgement it can see.
- 1.12.0 (2026-09-29): A CLIENT CAN DECLARE A SECTOR LIMIT, AND A LOOSER LIMIT IS SAID. The house limits (45% in one sector, 25% in one name) had one client override, the single-name limit, and it was honoured even when looser than the house one, in silence, against the code's own comment that a client may only tighten. Now profile_max_sector is declared beside profile_max_position, and signal_engine.concentration_limit is the one reader of both: a declared limit replaces the house one, tighter or looser. A looser one is a departure from the house view (CFA Standard III(C), Addressing Unsolicited Trading Requests, T0: discuss, explain, record - neither refuse nor silently accept), so while the book sits above the house level the report states it, once per sector or name: the weight, the house limit and the declared one (concentration_departures, printed by the report and the web Strategy page). Not as a per-holding HOLD: summarise_signals keeps one HOLD per holding, and a quality HOLD (strength up to 72) would have hidden it, or it theirs. The overweight test, the fund-first insertion, the contribution plan's sector line and the contribution dependency follow the declared limit; the fund-breadth rule stays at the house level because it describes an instrument, not a client. Reach on production 2026-09-29: 7 users, 5 with an advisor report, 3 whose largest sector is above 40% of holdings, 2 above 45%; 1 with a looser single-name limit already honoured in silence (50 against 25). 0 with a declared sector limit, because the field did not exist.
- 1.11.0 (2026-09-29): A DECLARED CASH RESERVE IS OUTSIDE THE ALLOCATION POLICY. The allocation lens added every declared cash item to the cash class, so a savings account and an emergency fund counted as money the policy could deploy, and a cash-over breach proposed moving money out of them. Liquidity needs are the first client constraint a policy is written under (CFA Standard III(C), Standards of Practice Handbook, T0). The client now declares a reserve in dollars (profile_cash_reserve, read by assetmix.load_reserve and by nothing else); assetmix.compute_mix removes min(reserve, cash) from the cash class AND from the base every percentage is a share of, so drift, the policy trade, the contribution plan, the Today card and the entitlement preview follow without re-deciding. A reserve larger than the cash on file is reported as a shortfall and is never made up from a GIC or a holding; the contribution plan rebuilds a short reserve before it invests anything. With NO reserve declared nothing moves: the value is unknown, not zero, and where a cash-over breach draws on savings the client entered the engine says how much of the cash that is and where to declare a reserve. Reach on production 2026-09-29, before code: 7 users, 2 with a declared policy, 1 with declared cash items, 1 whose leading action moved money out of them. Moves no published number until a client declares a reserve; 0 of 7 had one on the day it shipped, because the field did not exist.
- LIMITS OF THE CASH RESERVE. (a) It is DECLARED, never inferred: not from an item named Emergency, and not from emergency_fund_months, which cannot be turned into dollars without an expense figure nothing collects. A client who declares nothing gets the old arithmetic and a question. (b) It is taken from the cash class as a whole, idle brokerage cash and declared items alike; the engine does not say WHICH account holds it. (c) [Fixed in 1.12.0: the years-to-policy figure used to assume the reserve-rebuilding year repeated; only year 1 rebuilds it now.] (d) The retirement pool and the per-ticker rebalance planner do not read the reserve: neither includes declared net-worth cash. (e) CORRECTION to 1.10.0 limit (e): the 40 and 45 sector figures are two levels by design (stop adding at 40, trim above 45, signal_engine.SECTOR_CEILING_PCT), not three thresholds that should be one number.
- 1.10.1 (2026-09-29): CORRECTION - FIGURES PUBLISHED ABOUT THE OWNER'S BOOK IN 1.9.0 AND 1.10.0 WERE WRONG, AND THE ENGINE WAS NOT. They were taken from a book assembled by hand rather than from the engine: it priced holdings from the stock table only, which missed his fund holdings, and read cash from a table that does not hold it. It came out at C$72,038, 100% equity and 80.0% Technology. The engine's own drift block says C$122,094 - equity C$86,557 (70.9%), nothing in fixed income, and C$35,537 (29.1%) in cash-class assets - with Technology at 71.0% of holdings. So the published statements that a year of contributions reaches his policy, and that his breach was equity over target, were false: his breach is CASH over target by 19.1 points, the trade that closes it moves C$23,327 out of cash and requires no sale, and contributions alone take 4 years. The entries above are corrected in place. No shipped function changed: each reads the engine's blocks, so what the product computes was right throughout. What was wrong was what was SAID about one book, in this register, in three docstrings and in advice given to its owner.
- A FINDING THE CORRECTION EXPOSED, NOT YET BUILT: the C$35,537 of cash-class assets is two declared net-worth items - a savings account of C$25,500 and one named 'Emergency Saving' of C$10,000. The allocation lens counts both as cash available to the policy, so the rebalancing trade proposes moving C$23,327 'out of cash' on a book whose only other cash is C$37. An emergency reserve is a liquidity constraint, not an allocation to be optimised (CFA Standard III(C) requires the client's constraints to be considered before a recommendation). Nothing in the engine separates reserved cash from deployable cash, and the profile's emergency_fund_months is read by the prompt but by no rule. Reported for the owner to order; not fixed in this change.
- 1.10.0 (2026-09-29): THE REPORT SAYS WHERE THE NEXT CONTRIBUTION GOES. For a client still saving, what they are about to add outweighs what they hold: on the owner's plan 91.2% of the projected pot at retirement (C$1,612,222 of C$1,768,621, on the engine's own holdings total) is contributions not yet made. The engine already said how LONG contributions take to dilute a concentration (_compute_contribution_dependency); nothing said WHERE the money should go. The premise this work started from - that contributions were absent from the advisor - was wrong and is corrected here: the dilution horizon existed, the allocation did not. REACH, measured on production: 7 users -> 6 with holdings -> 1 with an annual contribution on file -> 1 with a declared policy as well. This is capability for one book today and is stated that way.
- THE CONTRIBUTION SPLIT IS AN IDENTITY, NOT A MODEL. After a contribution C the book is T + C; each class's target in dollars is its policy weight times that; what a class needs is its shortfall to that figure, never negative; C is shared in proportion to need, and a class already above target gets nothing. The needs sum to exactly C when no class is above its new target. NO RETURN IS ASSUMED AND NONE IS CLAIMED: growth is ignored, which makes every horizon conservative. Nothing is sold. On the owner's book as the engine measures it (C$122,094: equity 70.9%, cash-class 29.1%; policy 80/10/10, C$60,000 a year) the split is C$45,871 to equity, C$14,129 to fixed income and nothing to cash; contributions alone take 4 years to reach the policy, because cash sits above its target; Technology is 50.3% of the whole book and would be 33.8% if the equity money holds none of it, or 35.7% through the first screened fund. No external convention is borrowed because none is needed - the hand calculations in the test file are the check.
- THE TRADE AND THE CONTRIBUTION ARE TWO ROUTES TO ONE POLICY, and the report stated only the one that sells. The engine-enforced breach action already said 'direct new contributions here first' - in words, with no amount and no horizon. It now states the contribution route beside the trade in the same figures: where a year of contributions reaches the policy it says so, and where it would take longer it says how many years and that the trade closes the rest, rather than offering a slow route as an alternative.
- LIMITS OF THE CONTRIBUTION PLAN. (a) It needs a declared policy; with none it allocates nothing and asks for one. (b) It answers the ASSET-CLASS question and the sector question; it does NOT say which ACCOUNT the money goes into - registered room is reported elsewhere in the same report and is not re-derived here. (c) The contribution is assumed level at the figure on file. (d) Where the policy holds a class at 0% that the client owns, no horizon is stated, because adding money never removes a holding. (e) THREE THRESHOLDS NOW DESCRIBE 'CONCENTRATED' and they are not the same number: _OVERWEIGHT_SECTOR_PCT is 40, while _compute_contribution_dependency's target and the prompt's trigger are 45. This change uses the first and leaves the other two where they were; they should be one number, and which one is a decision for the owner.
- 1.9.0 (2026-09-29): THE FUND HALF OF THE DIVERSIFIER MENU IS NOW SCREENED AND MEASURED. Single stocks have always come from a screen over the scored universe. Funds came from the prompt line 'ETFs may come from your general knowledge' - the model's memory - while 245 scanned funds sat in etf_fundamentals that the candidate screen never queried. Measured on production before any code: 67 stored reports -> 49 carrying diversifier recommendations -> 245 recommendations -> 185 single stocks and 60 funds; of the 60, 23 were VXUS and 16 were VUG, which are the tickers the prompt itself used as examples, and VUG is 57.75% Technology by its stored sector weights while being recommended to diversify technology. The owner's own latest report was 5 single stocks of 5 on a book 71.0% Technology by the engine's look-through. The prompt now names no fund ticker at all, and a gate enumerates that as a class.
- THE FUND SCREEN REUSES DECIDERS THAT ALREADY EXISTED RATHER THAN CHOOSING NUMBERS. Sector weights must be known (unknown is not offered as a diversifier). Asset class is assetmix.classify_holding, because the stored data gives bond funds equity-style sector weights. The quality floor is the fund score's own top tier. And the 40% at which this report calls a client's book concentrated is the same 40% above which a fund cannot be offered - in the client's overweight sector, or in ANY sector, since a single-sector fund swaps one concentration for another. On the live universe of 2026-09-29: 245 funds -> 63 with no sector weights -> 7 not equity -> 112 below the quality floor -> 14 concentrated in the overweight sector -> 7 narrow -> 42 eligible.
- TWO DESIGNS WERE MEASURED ON LIVE DATA AND DISCARDED, and both are recorded so neither is rebuilt. Ranking by the book's post-add effective number of sectors SATURATED: the top 25 funds all scored 1.80-1.81, so the order was noise and was led by covered-call utility funds. Ranking by weight in the overweight sector ALONE put XLE - 100% Energy - first, and the engine's own text called it a broad fund. The breadth rule above is what that second failure produced.
- _SAME_EXPOSURE_L1_PP = 5.0 is a DECISION, taken from a measurement and not fitted to any outcome. It shapes only which screened funds the prompt is shown. Trackers of one index sit 0.1-1.7 points apart in summed sector weight (IVV/VOO 0.1, XIC.TO/ZCN.TO 0.8, TTP.TO/XIC.TO 1.7); funds that are genuinely different sit at a median of 53.5 across all 861 eligible pairs. Without it the first broad menu was five trackers of the Canadian composite. The short menu is also split across account currencies, because the report tells the model to deploy an account's cash without converting it.
- WHERE A SECTOR EXCEEDS 40% AND NO SCREENED FUND WAS RECOMMENDED, THE ENGINE PLACES ONE FIRST. It is inserted after the model has written, because a prompt instruction is a request; it carries NO position size, because none was computed; and the model's own picks are kept behind it. Every fund recommendation, inserted or chosen, states its measured weight in the overweight sector, or that it was not measured, or that it is itself concentrated there and does not diversify it.
- LIMITS OF THE FUND SCREEN, stated rather than left to be found. (a) Coverage: sector weights are on file for 182 of 245 funds, so 63 cannot be offered however suitable. (b) The quality floor inherits the fund score's trailing-return component - 35 of 100 points, 50 for a growth objective - so the menu leans toward funds that have done well; it is a floor and not the ranking key, but it is not neutral. (c) The measure is HOLDINGS-BASED: it says what a fund owns, not how it moves with the book. Only 13 funds carry price history, so the return-based effective-bets measurement used for single stocks cannot run for them. (d) It addresses SECTOR concentration only: bond and cash funds are excluded by design, and the asset-class policy breach is the allocation lens's job. (e) NO FORWARD CLAIM is made - no return is stated or implied for any fund, and nothing here is validated (CFA Standard V.A).
- 1.8.0 (2026-09-26): THE ADVISOR NOW READS THE RETIREMENT INCOME PLAN, AND THE CLIENT CAN SWITCH IT OFF. advisor_engine imported retirement_sim (the goal Monte Carlo) and NEVER decumulation - zero references in the module. So the report could reason about reaching a target BALANCE and had nothing to say about the INCOME that balance must produce: sustainable spend, CPP/OAS timing, the RRIF minimum that forces withdrawals from 71, and the OAS recovery tax those withdrawals can trigger. The goal lens was answering a different question and standing in for this one. The lens reads api.retirement_api.plan_for_user - the SAME assembly GET /retirement/plan serves - so a retirement figure in the narrative cannot contradict the Retirement panel the client is reading; the route's own copy of the steps was extracted into that function and into retirement_ages in the same push, and a gate fails the build if advisor_engine ever calls build_plan directly.
- REACH, measured on production 2026-09-25 BEFORE building: 7 users -> 2 with birth_year (the plan cannot be dated without it) -> 2 who also have retirement_age -> 0 with ret_spend_target -> 5 who have ever generated an advisor report. HARM SET: 2 users have a plan AND read advisor reports. This adds a lens for 2 of 7 and MOVES NO PUBLISHED NUMBER - decumulation.py is byte-identical to the deployed commit and no existing payload key changes value.
- NOBODY HAS SET A SPENDING TARGET - 0 of 7 - so the gap analysis is dark for every user, and that fact shaped the design rather than being a footnote. With no target there is NO gap: gap_cad is null and never 0, spend_target is listed in missing_inputs, the prompt is forbidden to assume a spending level or to call anyone on track or behind, and both renderers say the target is absent instead of showing a verdict. A default would be the model telling a client they are on track against a number nobody chose (CLAUDE.md 9, unknown is never a number). Naming the missing input is also the only mechanism that ever gets it filled in.
- THE CLIENT CAN OPT OUT (owner request, 2026-09-26), and the switch is a DEDICATED setting rather than 'leave birth_year blank'. birth_year is in MINIMUM_SUITABILITY_FIELDS and sets years_to_retirement, which drives risk capacity for the WHOLE report - using it as the off switch would silently degrade every other piece of advice to buy silence on one lens. ret_in_advisory is read in exactly one place (advisory_includes_retirement), shared with the save route so a form that omits it cannot switch advice back on; it is checked BEFORE anything else, so an opt-out never surfaces as 'we need your birth year'; and it defaults to ON for everyone who predates it. It suppresses the ADVISORY lens only - GET /retirement/plan ignores it, because the panel is where the switch lives.
- THE PROJECTION IS NOT VALIDATED (CFA V.A). It is a projection under stated assumptions - pre-tax, in today's dollars, net of the book's fees when they are known (2026-10-06; the 2026 FP Canada Projection Assumption Guidelines state that management fees must be subtracted to obtain the net return) and labelled gross when not. It is not validated, not backtested and not predictive, and no part of the lens, the prompt or either renderer may describe it as such; a gate asserts those three words appear near the lens only inside the sentence that denies them.
- 1.7.1 (2026-09-25): THE DISCLOSURE'S OWN REASON WAS RESTATED, BECAUSE IT HAD BECOME FALSE. The Growth-pillar note said the pillar's sign was untestable because price history covered 88 of 1,023 scored names. The bounded universe backfill closed that gap (2,786 of 2,792 scored names; 988 of the 1,059 scored on 2026-06-08), so the remaining blocker is the 126-day window closing 2026-10-12, not absent data. A disclosure that keeps citing a resolved obstacle overstates the limitation as surely as omitting it would understate it.
- 1.2.0 (2026-09-24): TWO DEPENDENCES THE RECOMMENDATION RESTS ON ARE NOW DISCLOSED BY THE ENGINE RATHER THAN ASKED OF THE MODEL (CFA Standard V.B). (a) GROWTH-PILLAR DEPENDENCE: a recommended single name whose Deep Score owes the Growth pillar more than an evenly-scoring name would (above PILLAR_MAX['Growth'] / 100 = 25%) carries that share, in its own `why`, with the two findings that bear on it. Measured on the production scan of 2026-09-24 over 2,577 scored names: the Spearman between the shipped score and the score with the pillar removed is 0.9700, but 45 of the top 50 names are in that set only because of the pillar and 6 of the 8 the live screen surfaced were Growth-lifted, mean 20.6 of 25 against a universe mean of 12.5. (b) SINGLE-NAME CONCENTRATION: the share of the book held as individual securities rather than funds, against the 20% default budget, with its sources. NEITHER CHANGES A SCORE, A RANK, A PICK OR A SIZE - the same names are recommended in the same order at the same weights; what changed is that the report says what they rest on.
- The Growth pillar is disclosed, NOT re-weighted. Until 2026-09-25 the reason was missing data: price_history covered 127 tickers of which 53 were held, so a forward-return test was n=88 and selection-biased toward the book. The bounded universe backfill closed that gap — 2,786 of the 2,792 scored names now carry daily closes, and 988 of the 1,059 names scored on 2026-06-08 span that scan to today. The remaining blocker is TIME: that cohort's 126-day forward window closes 2026-10-12. Until it is graded, re-weighting the composite on the literature alone would substitute one unvalidated prior for another.
- The single-name budget is DISCLOSED, NOT ENFORCED. No trade is blocked, no recommendation is resized and no plan is rejected for breaching it; the report states the figure and says so explicitly.
- 1.3.0 (2026-09-24): A POLICY BREACH NOW NAMES BOTH SIDES OF ONE TRADE, IN DOLLARS, WITH WHAT FUNDS IT. _enforce_policy_first led with breaches[0] - the single worst class - and stated the action in percentage points with no amount, no funding source and no tax consequence. On a book 24.2pp over in cash and 20.0pp under in fixed income the client was told 'Reduce cash toward 5%' and never told where the money goes; the destination WAS the second breach. Measured on production 2026-09-24: 6 users carry settings, 2 have a declared policy, and 2 of those 2 are in breach - BOTH holding 0% fixed income against a 20% target while over on another class (neha equity 100.0% vs 75%, vishal cash 29.2% vs 5%). Every class now moves to its declared target so the legs balance by construction, and the funding split is reported: vishal's C$242,000 rebalance is funded entirely from cash and realises nothing, while neha's C$109,250 requires selling equity.
- THE TAX ON A REBALANCE IS NOT ESTIMATED, only the funding split. Which lots in which accounts are sold decides it, and _policy_rebalance_trade is handed neither, so it reports how much requires a disposition and stops there rather than inventing a plausible figure. Money the look-through could not classify is excluded from the trade and named, never absorbed into a class.
- VANGUARD'S 200bp-TRIGGER / 175bp-DESTINATION REBALANCING BAND WAS CONSIDERED AND DECLINED (CLAUDE.md 14, cargo-cult test). It is derived on target-date funds rebalanced at scale, where the question is when to trim a portfolio NEAR its targets. That precondition fails here: the declared band is 5pp and the measured breaches are 20-25pp, four to five times outside it, so tuning the trigger would move nothing on any live book. The band remains the one the client declared, and every leg targets the declared policy rather than a band edge.
- 1.4.0 (2026-09-24): CAPITAL-GAIN STAGING WAS COMPLETE, CORRECT AND DEAD. _annual_gain_budget_cad reads investor_profile['annual_gain_budget_cad'] and _gain_staging_note turns it into advice; both were tested. The key appeared in exactly two places in the repository - its reader and its reader's test - with NO writer in src/, web/, ios/, android/, JSON or HTML, so the note had never fired on any book and could not have. Three separate breaks, of which the saved plan named one: (a) no KYC field, so no client could state a budget; (b) no mapping from the stored setting profile_annual_gain_budget into the profile key the engine reads; (c) TWO of the four live callers of apply_tax_and_allocation passed no investor_profile - api/portfolio._bundle (web + mobile) and intel_agent (the AI chat) - so even a stated budget would have reached the Strategy report and not the other three surfaces showing the same signals. All three are closed and the call sites are now enforced as a CLASS by walking the call graph, so a new surface must pass the profile or fail the build.
- GAIN STAGING CHANGES NO PUBLISHED NUMBER TODAY, and that is the honest statement of what it is: measured on production 2026-09-24, 6 users -> 4 carry a positive unrealised gain in a capital-gains-taxable account (the registry's own capital_gains_taxable field decides which; C$19,300 / C$99,695 / C$24,756 / C$3,389,870) -> 0 have a budget on file. Until a client states one, _annual_gain_budget_cad returns None and stages nothing by design - an absent budget is UNKNOWN, never zero and never unlimited. This ships capability, not a recovered loss.
- The staging advice rests on CRA's own T4037 Capital Gains guide (2025 ed., retrieved 2026-09-24), which states a capital gain is reported in the calendar year of the disposition - which is what makes 'split it across two tax years' a real instruction rather than folklore. The cargo-cult precondition HOLDS here, unlike the rebalancing band declined in 1.3.0: four live books carry unrealised gains that a single-year realization would treat very differently from two.
- 1.5.0 (2026-09-24): THE PROPOSED POLICY NOW REACHES THE PANEL THAT ACCEPTS IT. The report told clients 'your questionnaire proposed X - accept it in Investment Policy', and GET /assetmix, which feeds that panel, did not carry X: the client arrived at an empty form defaulted to a house 80/15/5 and had to retype three numbers from memory. The route now returns `proposed` from the SAME decider the report uses (_proposed_policy_from_questionnaire), only while no policy is on file, and the panel prefills the editor from it - PREFILLED, NEVER APPLIED, so the policy on file is still one the client deliberately saved (CFA III.C). Measured on production 2026-09-24: 6 users -> 2 with a declared policy -> 1 with a proposal no surface could act on (holding 100% equity against a proposed Conservative 45/50/5) -> 3 who have answered nothing. assetmix is absent from contract.py and has no iOS/Android consumer, so the frozen mobile clients are untouched.
- THE ALLOCATION LENS STILL DOES NOT REACH MOST CLIENTS, and that is an adoption gap rather than an engine gap: of 6 users with settings on 2026-09-24, 2 have a declared policy, 1 (demouser) has a Conservative 45/50/5 mix PROPOSED by the questionnaire that has never been accepted while holding 100% equity, and 3 have answered no questionnaire at all. GET /assetmix does not carry the proposal, so the panel where a client would accept it cannot show it.
Entry points & executable coverage
Superficial-loss determination (ITA s.54 / s.251.1)
superficial_loss · owner vishalsinha · version 1.0.0 · last reviewed 2026-09-24 · portfolio_db.py
Monitoring
MEASURED 2026-09-24 - 0 of 4 taxable loss positions disagree; 0 declared spouses
Metric: harvest candidates where the Tax Centre and the signal engine disagree. Grader: manual prod probe (scripts/reach.py + the harvest funnel). Log: none.
Validation
invariant tests/invariants/test_harvest_superficial_one_decider.py, test_affiliated_person_invariants.py, tests/test_superficial_loss*.py - still-held, split rows, affiliated scope, settlement date, and failure withholding the all-clear; 3 entries in tests/parity/test_gate_injections.py prove them red
Assumptions
Limitations
- 1.0.0 (2026-09-24): REGISTERED BECAUSE IT MAKES A TAX DETERMINATION THE CLIENT ACTS ON, and because it briefly had two implementations. portfolio_core.build_tax_data ran its own scan for harvest candidates - 'any BUY of this ticker in the last 30 days' - beside the rigorous quantity test here. The local scan had no quantity or shares-still-held test, no affiliated-person scope, counted $0 split rows as reacquisitions and used trade rather than settlement date. Both directions were wrong: it blocked a harvestable loss when a rebuy had since been sold, and it advertised 'Harvest now' on a loss a spouse's rebuy denies. Measured on production 2026-09-24 before the change: 6 users -> 4 taxable loss positions -> 0 where the two answers differed, and 0 users with a declared spouse, so no published number moved. One decider now, enforced by tests/invariants/test_harvest_superficial_one_decider.py.
- MODELS ONLY THE INDIVIDUAL/SPOUSE LIMB of affiliation. s.251.1 also affiliates corporations a person controls and certain partnerships and trusts; this product has no data on those, so they are out of scope rather than handled. Affiliation is honoured only when BOTH partners declare it, because acting on it denies one of them a deduction because of the other's trade.
- The harvest check asks the rule PROSPECTIVELY ('if this were sold today'), so the forward half of the +/-30-day window is necessarily empty - a rebuy AFTER the sale would deny the loss and cannot be seen in advance. The client is told to hold through the window rather than given a guarantee.
- No dollar tax figure is attached to a denial on this path: the amount depends on the client's marginal rate and the year's other dispositions, which this primitive is not given.
Entry points & executable coverage
Retirement Monte Carlo
retirement_monte_carlo · owner vishalsinha · version 1.5.0 · last reviewed 2026-10-06 · retirement_sim.py
Monitoring
NOT MONITORED
Metric: realised path vs simulated distribution. Grader: none. Log: none.
Validation
process invariant tests/invariants/test_retirement_mc_invariants.py — the drawn process must deliver the compound return the plan states, the conversion must be disclosed, and sigma=0 must still reduce exactly to the deterministic path; proven red against the pre-fix behaviour
Assumptions
Limitations
- 1.5.0 (2026-10-06): THE ODDS RUN NET OF FEES, TO THE CLIENT'S HORIZON, WITH SENSITIVITY ROWS, AND THE ADVISOR QUOTES THEM. api.retirement_api.simulate_plan is now the ONE simulation (the route and the Strategy Advisor's retirement lens), so the report no longer prints a median-case spend without the spend that lasts in 90% of paths. Every pass runs on the plan's own rate (net of the book's fund MER and any declared advisory fee, via return_basis; gross and labelled when an MER is unknown) and to the plan's FP Canada 25%-survival horizon. Four extra passes on the same seeded draws give the target's odds with the horizon five years either side (PAG 4(e)) and the return one point either side (PAG section 2 scenario testing) — measured at ~10-20 ms a pass against ~170 ms for the 90% spend search already in the block.
- 1.4.0 (2026-09-30): THE ODDS RUN ON THE POT PROJECTED AT RETIREMENT. The success odds and the year-by-year path started from today's pot while the sustainable spend beside them was computed on the pot projected at retirement; both now start from the projected pot. The sustainable spend is the MIDDLE case and under this simulation it lasts in roughly half of paths, so with a spending target set the panel also shows that spend's own odds and the spend that lasts in 90% of simulations (spend_at_success, solved on the same seeded draws solved to within C$250 and rounded down, so it errs on the safe side). success_pct is None, never a figure, when there is no horizon (a retirement age at or after 95). The simulation block now renders: a string-format error in its assumptions note had suppressed it, so no odds were shown until this fix (audit 2026-09-30).
- 1.3.0 (2026-09-25): AN UNCLASSIFIABLE BOOK PUBLISHED A CERTAINTY. blended_sigma normalised its weights with max(1e-9, total), so a mix of (0,0,0) returned 0.0 volatility rather than unknown. At sigma 0 the Monte Carlo is exactly the closed-form future value - one path, p10 == p50 == p90, and a success probability of exactly 0.0 or 100.0 - so a portfolio nobody could describe was shown a 100% probability of success, and implied_real_return likewise reported 0.0% as though it were a measurement. Normalising a PARTIAL mix is right and the callers rely on it; normalising an ABSENT one is the invariant-over-rows-present failure, and the published figure is what CLAUDE.md 9 names outright: unknown is never a number. Both helpers now return None, the simulator abstains, and the advisor no longer coerces a missing sigma back to 0.0. ABSENT IS NOT ZERO: sigma == 0.0 passed deliberately still reduces to the deterministic answer, pinned to the dollar elsewhere. MEASURED REACH ON PRODUCTION BEFORE THE FIX: 6 users with settings -> 2 who see a retirement plan at all -> 0 whose book classifies to sigma 0 (both resolve to 100% equity, sigma 0.17). A live mechanism with an EMPTY harm set, fixed because it is cheap and enforces a standing rule, not because anyone was harmed.
- Draws i.i.d. normal annual returns (audit E2): no fat tails, no serial correlation, no regime structure. MEASURED 2026-09-10 rather than left as an aspiration, both upgrades rescaled to the same standard deviation so only the shape changed: fat tails (Student-t) move success by +0.3pp at df=5 and -1.0pp at df=8 — less than this simulation's own seed-to-seed noise of 0.99pp — because over a 30-year horizon the variance dominates and the tail shape washes out. Mean reversion (AR(1)) moves it +1.2pp at rho=-0.15 and +3.0pp at rho=-0.30, which is material but OPTIMISTIC and rests on an empirical claim whose sign is contested for annual equity returns. Neither adopted: one would add an ungrounded parameter for no measurable effect, the other would flatter every plan on a contested premise. The limitation is disclosed in the assumptions note the user reads.
- Asset-class volatilities are unsourced point estimates, in a model that reports a probability of success.
- Sequence-of-returns risk is represented only through the i.i.d. draw.
- FIXED 2026-09-10: the stated COMPOUND real return was fed in as the arithmetic mean of the draw, so the simulation delivered ~sigma^2/2 less than the plan displayed — measured at 1.45%/yr against a stated 3.00% for an all-equity mix, while simulate_deterministic compounded the same input and got exactly 3.0%. Success probability was understated by 12.0 to 15.9 points depending on mix. mu is now geometric + sigma^2/2 and both figures are disclosed in the payload.
- The PAG adds ~0.5% on the equity portion when converting its own published figures. That term is specific to their derivation and is NOT applied here — an open decision worth about 3 further points of success probability.
- ASYMMETRY, now DISCLOSED rather than silent (2026-09-10): volatility is blended from the user's ACTUAL asset mix while the return assumption is a single mix-blind number, so an equity-heavy user was charged equity volatility against a balanced-portfolio return. implied_real_return() now reports what the mix implies on the 2026 PAG — 2.99% at 60/40, 3.42% at 75/20/5, 4.25% all-equity — beside the plan's own figure, in the payload. The plan still RUNS on the user's number: adopting the mix-implied figure as the default would make equity-heavy plans more optimistic, which is the direction where being wrong costs the user most, and would additionally need a fee assumption because PAG figures are gross of fees. That adoption is an open decision for the owner.
Entry points & executable coverage
Risk questionnaire → suggested mix
risk_profile_questionnaire · owner vishalsinha · version 1.0.0 · last reviewed 2026-09-30 · riskprofile.py
Monitoring
NOT MONITORED
Metric: suggested mix vs the mix the client adopts. Grader: none. Log: none.
Validation
property test tests/test_riskprofile.py — every profile's targets sum to 100 over all 1,024 answer combinations, and the horizon and capitulation caps are applied and named
Assumptions
Limitations
- Five questions, scored 0-3 each into five even bands; the output is a SUGGESTION for the investment-policy editor and never applies itself.
- The mixes and the band edges are house judgement calls with no recorded external basis.
- Willingness is read from self-reported answers only; capacity is read from the horizon answer alone, not from the client's income, assets or obligations.
Entry points & executable coverage
Historical crisis replay + expected shortfall
historical_stress_replay · owner vishalsinha · version 1.0.0 · last reviewed 2026-09-30 · stress_replay.py
Monitoring
NOT MONITORED
Metric: none — a replay of history has no forward outcome to grade. Grader: none. Log: none.
Validation
invariant tests/invariants/test_stress_replay_invariants.py — today's weights on actual closes, coverage stated, expected shortfall <= VaR <= mean, dollars on the covered value only
Assumptions
Limitations
- Replays TODAY's weights on actual closes: what the current book would have done in each window, not what any book the client held did.
- Holdings without closes in a window are left out and the coverage is stated; three windows only, all since 2007.
- Expected shortfall is a 1-day, 95% figure over the last 756 common trading days of the covered holdings; it describes the past distribution, not a forecast.
Entry points & executable coverage
Beta-adjusted downside table (Portfolio page)
beta_stress_table · owner vishalsinha · version 1.1.0 · last reviewed 2026-10-06 · advisor_engine.py
Monitoring
NOT MONITORED
Metric: none — the shocks are illustrative, not predictions. Grader: none. Log: none.
Validation
invariant tests/invariants/test_beta_stress_invariants.py — cash is never shocked, a holding with no beta is excluded and named (never given one), the report rows use the same decider, a covered CDR reads its underlying's beta, the clamp holds
Assumptions
Limitations
- Linear in beta: a holding falls beta x the index shock, beta clamped to [0, 3] and each holding floored at zero; no correlation, liquidity or volatility regime is modelled.
- 1.1.0 (2026-10-06): a holding with no beta is EXCLUDED from the shock and named with its C$ value; the loss and loss % are over the modelled holdings and 'Portfolio after' is not stated while any holding is unmodelled (1.0.0 gave it the covered holdings' average beta, so a GIC beside high-beta stock fell harder than the market). With no beta anywhere nothing is modelled: loss, value after and vs-flat are all not stated (labelled 'no betas on file — not stated'; until 2026-10-07 it showed a flat index drop, a beta of 1 for every holding). The Strategy report's beta rows use the same decider (_stress_population).
- Side effect of the CDR look-through, measured on a test book (tests/invariants/test_cdr_beta_side_effects.py): late_cycle_trim and deteriorating_outlook_trim now fire on a covered CDR in a registered account; a taxable winner still stands down; the contraction SELL still needs a Deep Score a CDR lacks; an ADD of a covered CDR is now judged against a declared drawdown limit.
- An all-cash book is labelled cash_only: nothing was beta-adjusted.
- A CIBC CAD-hedged CDR on a US company takes its US underlying's beta (cdr.BETA_FROM_UNDERLYING; CIBC: 'your returns depend on the performance of the underlying shares — regardless of currency fluctuations', T2, retrieved 2026-10-06). Other receipts are not covered and stay unmodelled: a receipt that is not hedged, or is on a non-US company, would not carry the US listing's beta.
- Cash is carried unshocked. Beta is the vendor's trailing figure and is not re-estimated; where the vendor has none the scanner computes one from two years of daily returns, and failing that it currently stores 1.0 (fundamentals_metrics) — such a holding is modelled at a beta nobody measured, and this table cannot tell it apart (open, reported 2026-10-06).
Entry points & executable coverage