1. What the numbers mean
The application is a research monitor. Its score anchors, weights and alerts are uncalibrated. The headline Stress level is the underlying Stress Index multiplied by 100 / 45: 100 marks severe stress and is not a maximum. Component and scenario scores retain their own 0–100 scales. Neither scale is a probability. The app estimates exploratory historical shares for the first time its own index reaches a threshold; it does not supply a confident collapse date or a calibrated probability of an externally defined crisis.
Missing or stale observations do not become zeros. A page with poor coverage does not imply low risk. Unknown institutional exposures do not imply no connections. The supplied dated baseline is never treated as live input.
2. Observation and provenance model
Each observation has an ID, metric ID, measurement date, value, exact unit, named source, source URL, provenance type, quality, retrieval time, knowledge time, vintage-verification flag, notes and demo marker. A changed source reading appends a revision; unchanged readings do not create artificial history. Downloads use collection time as known_at, even when measurement dates are historical. This avoids pretending that a current-vintage historical download was known in the past.
Imported vintage flags are user attestations, not certification. The software cannot establish that a provider actually published a historical reading on a claimed date. It cannot reconstruct unavailable vintages.
For the same metric/date/source, the latest revision is used. Among different sources for the same metric/date, selection favors declared provenance trust times data quality, then recency. Conflicts remain visible in source variants. A disagreement diagnostic scales the range of conflicting values against the larger of 3% of their median level, 5% of the heuristic normal-to-severe span, or 0.01. It is capped at 1 and reduces confidence by up to 50%.
This comparison is only like-for-like metric IDs/dates/units. It does not automatically reconcile payrolls with job postings, CPI with commodity prices, or other different economic concepts. Correct series definitions remain the analyst's responsibility. Historical derived series inherit the weakest parent provenance and quality rather than laundering unverified imports into trusted market data.
Default provenance weights are editable. Government-origin direct transaction/operational data can be high-trust; official estimates are lower-trust by default. Neither official status nor private ownership proves truth or falsehood. A declared high-quality source can still be wrong.
3. Standardization and changes
For a metric with sufficient history, the robust score compares the latest reading with prior observations, excluding the latest observation from its reference set:
robust_z = (current - median(previous)) / (1.4826 * MAD(previous))The prior window uses up to 252 daily, 52 weekly, 12 monthly or 4 quarterly observations. At least 20 prior readings are required. A zero MAD falls back to population standard deviation; a completely constant reference window produces 0 for an unchanged reading and unavailable for a different reading rather than an invented infinite score. Z-scores are limited to [-8, 8].
Percentiles use up to five years of the observed frequency and report the actual available span. The app does not call six months of data a five-year history. Changes are 1, 5 and 20 observations, not necessarily calendar days. Acceleration is change_5 / 5 - change_20 / 20. Daily derived returns spanning implausibly large gaps are omitted.
Derived SOFR–IORB and yield-curve spreads require matching observation dates. A five-observation yen appreciation is (previous USDJPY / current USDJPY - 1) * 100, so yen strengthening has the correct positive sign. Yield changes are in basis points, not percent.
Two derived series give the design's "cause-sensitive yield interpretation" a concrete form. The sell-America signature is the 5-observation 10-year yield change on days when the dollar fell over the same observations (the broad trade-weighted index where it exists, the same-day ECB-based DXY on the days the weekly H.10 release has not reached), and zero otherwise; yields rising with a falling dollar is the foreign-confidence and basis-unwind pattern of April 2025, whereas yields rising with a firm dollar (2013, 2022) is not. Its one-year counterpart applies the same rule to the 30-year yield over 250 observations and enters as a structural vulnerability. Both depend on the Federal Reserve's dollar index, which publishes with a lag of up to a week, so they carry a 14-day freshness limit. The foreign share of Treasury debt, foreign holdings divided by total public debt, is a quarterly structural input that lags close to a year with its Treasury Bulletin source; a year-old reading is still used because the quantity moves slowly.
4. Metric and scenario scores
A metric's level component maps its explicit normal and severe anchors linearly into [0,100], including reversed-risk measures such as debt-service coverage. In the default blend mode, an available robust z-score produces 75% anchored level plus 25% direction-oriented positive deviation (scaled over three z units); without it, the anchor score is used. Metrics marked score_mode='level' use the anchored level alone. ENSO inputs first take the absolute anomaly, so both warm and cold phases can contribute vulnerability while their displayed observations retain their signs. All anchors and transforms are inspectable in crisis/catalog.py.
Climate, reservoir, drought, observed monsoon, food-price and health inputs use level-only scores. A small departure from a stable short record therefore does not receive an extra rolling-anomaly boost. These physical and health anchors also remain fixed when active-stress financial anchors are referenced to 2008/2020. This is a choice of monitoring scale, not validation of an economic damage model.
Stale readings are excluded from scores. Freshness thresholds depend on the series frequency and release cadence. Effective quality is provenance trust times declared quality, excluded after the freshness limit and reduced for like-for-like source conflict.
Metrics are quality-weighted within explicitly named groups before those groups contribute. Vulnerability V averages group scores using mean effective quality per group. The headline active-stress component instead gives each available group one share of its intensity and breadth calculations, as specified in section 11. Daily, weekly, monthly and quarterly V/S inputs are eligible when fresh; monthly defaults and food prices are not discarded for being slower observations. Grouping limits duplicate votes, but remaining correlations across groups mean they are not independent economic risks. Coverage measures observed quality against all configured relevant metrics, including unavailable series.
Reviewed news and current social aggregates can adjust an existing scenario stress score; neither can create one from entirely missing measurements. They do not directly enter the combined headline or its timing profile. Scenario index:
index = 0.35 * V + 0.45 * S + 0.20 * V * S / 100These are prototype weights from the user's design, not a fitted final model. Both V and S are required. The separate rated-scenario composite averages eligible A–F scenarios only when at least three have scores and at least 45% coverage. This coverage-limited composite is distinct from the headline Stress level described in section 11. Inflation/fiscal context M is not a seventh independent crisis probability. The prototype does not yet estimate a learned nonlinear fiscal multiplier or separate cyber-by-funding interaction coefficient.
A cross-market confirmation count in the spirit of the design's Layer D is also reported: the number of distinct stress measurement groups scoring at least 60 with confidence at least 0.6 across all scenarios, with the group names. It is descriptive; alerts still follow the scenario rules above.
When a scenario or the overall index is unrated, the state carries a rating_gap: the fewest additional catalog inputs that would lift the scenario past the coverage floor and supply any missing V or S layer, chosen greedily by declared provenance weight. The projection assumes each suggested reading arrives fresh at full declared quality, and stale readings count as missing. It is coverage arithmetic that names the gap; it is not a judgment about which markets matter most, and it never changes a score.
5. Evidence and social safeguards
Automatic RSS parsing retains linked headline metadata and tentative keyword classification. The item is pending and contributes nothing until reviewed. SEC collection similarly identifies filing metadata, not ratios or loan impairments. Analyst review must check primary documents, entity, event date, direction, severity, scenario, directness, systemic link and independent corroboration.
Canonical links remove common tracking parameters. Items sharing a cluster contribute only the strongest positive and strongest negative evidence within that cluster. Cross-domain syndicated or semantically duplicate stories must be manually assigned the same cluster. URL clustering is not a complete semantic deduplicator. A claim and its rebuttal can coexist.
An event's strength decays exponentially with type-specific half-life (for example, gates 45 days, defaults 30, outages 3). Only reviewed items from the preceding year enter current event scoring. Net news contribution is bounded and symmetric:
news_points = cap * (tanh(sum_positive_strength / 2) - tanh(sum_negative_strength / 2))Default cap is 20 points. Source credibility, directness and analyst-attested independent confirmation affect strength. Social-origin events cannot be converted to primary factual confirmation merely by reviewing them.
Optional Reddit/Hacker News sampling creates only aggregate counts and keyword-based distress/fear features. Authors and source links are deduplicated transiently, then discarded. Raw post text, Reddit links and author identifiers are not stored. Aggregates are retained 45 days; separate scenario scoring uses samples less than 48 hours old. The default scenario social adjustment is capped at 5 points, configurable up to 10; it is not a headline contribution. Sampling is partial, excludes Reddit comments, and is not a validated sentiment model. It does not establish causal evidence or a panic rate in the population. No full narrative-diffusion model or first-hand-account classifier is implemented.
6. Alerts and regimes
GREEN/YELLOW/ORANGE are heuristic monitoring labels, not formal crisis-state classifications. An overall alert needs adequate evidence. RED additionally requires multiple stressed scenarios, multiple funding-related groups and a reviewed direct systemic funding event. BLACK additionally requires separately clustered, independently confirmed direct credit-contraction observations and evidence of failed stabilization.
High yields, falling stocks or social alarm alone cannot trigger the severe systemic labels. Absence of an alert is not proof of safety, especially with missing feeds. A contractual redemption cap is not automatically a gate, insolvency or default; the analyst must distinguish the facts.
Regime descriptions are explicit rule-based hypotheses, not fitted Hidden Markov probabilities. Policy and capital buffers remain accepted calibration features; their effectiveness is not quantitatively estimated in the live monitor. Context records that do not feed the headline are hidden from its dashboard input views.
7. Exposure graph, snapshots and explanations
Stored exposure edges require institutions, relationship, amount, unit, date, confidence and source. These records describe an observed subgraph, not a complete counterparty model or a calibrated contagion score. They do not enter the combined headline; the unused network dashboard view is hidden. Cross-unit amounts are not summed, and existing source records are retained.
Snapshots are append-only through the application and include model settings with secret values removed, selected observation IDs, evidence, explanations, versions and a SHA-256 integrity hash. Stored snapshots omit the 90-point plot arrays, which are display data reproducible from the stored observations; every scored value, source and evidence reference is kept. Each snapshot also writes a small summary row (overall index, coverage, scenario indices) that the history chart and session deltas read, so neither reparses full payloads and busy days can no longer crowd earlier sessions out of the chart. Reviews record before/after data in an audit table. Hashes can reveal accidental changes; the database is not cryptographically signed or protected against a malicious local administrator. Delta panels compare existing stored sessions; the app never fabricates prior dashboard history.
Details reports additive displayed-index-point contributions from the actual included factors. The interaction 0.20 × V × S / 100 is split equally: V receives 0.30 × V + 0.10 × V × S / 100, and S receives 0.50 × S + 0.10 × V × S / 100. Each budget is distributed using the factors' actual quality-weighted group shares; S shares include intensity and elevated-group breadth. Reference normalization, clipping, index rounding and the display conversion are reconciled so factor and group allocations sum to the displayed headline. If V is unavailable, the entire fallback score is allocated to S. An included zero-score factor can have zero allocated points.
This partition is neither a marginal sensitivity nor a causal impact: removing a factor changes group denominators, breadth and possibly the reference calibration, so the resulting score change need not equal its allocated contribution. Raw upstream observations identify the derived factors they support and receive no duplicate points. Ranked diagnostic signals and scenario deltas remain distinct from this additive breakdown.
What-if observations exist only for the request and do not alter the database. Derived series—including yield changes, curve and repo spreads, yen/equity moves and FAO annual price changes—are recomputed from overridden raw inputs, while an explicit value entered for a derived metric takes precedence. Changing weights changes interpretation and should be documented.
8. Calibration research tool
calibrate.py accepts externally prepared, point-in-time daily feature rows and externally defined systemic onset labels. The target is onset during the next calendar day, conditional on not already being in a crisis. The caller must supply all at-risk days and meaningful episodes; this tool does not derive labels from price crashes.
Input schema is in templates/calibration_template.csv. Features are 0–100 vulnerability, active stress, contagion, confirmation, policy buffer and capital buffer. Knowledge time must not exceed decision time. Outcome knowledge must follow the full next-day window and cannot be future-dated. Unverified-vintage rows are refused, and ongoing-crisis rows are excluded. User attestation is still not independent verification.
The research split is chronological: first 75% train, last 25% hold out. Training labels must have matured before the holdout starts. Independent episode IDs cannot cross the split; only one positive onset per episode is accepted. Guardrails require at least 500 at-risk observations, five training onsets and two held-out onsets. These are refusal guardrails, not a claim of sufficient sample size or reliable validation.
The fit is ridge-regularized logistic daily hazard (maximum a posteriori under a Gaussian penalty), with a Laplace inverse-Hessian covariance approximation. Diagnostics include Brier score, baseline skill, five reliability bins and threshold-dependent precision/recall. It is not a full Bayesian hierarchical six-pathway model or robust uncertainty analysis.
Every output says RESEARCH_ONLY_NOT_APPROVED_FOR_FORECASTING and forecast_activation: false. No artifact can activate dashboard forecasts. Future factor dynamics, cross-scenario dependence, intervention effects, external validation, regime transportability and uncertainty in labels still need development and review before meaningful multi-month onset distributions can be produced.
9. Access and security scope
The server binds to 127.0.0.1, validates Host and same-origin writes, requires a per-process request token, and serves no third-party scripts. Outbound collectors validate public HTTPS endpoints, verify TLS and use time/size bounds. No access-control bypass, paywall bypass or scraping of blocked social endpoints is included.
This is a single-user research tool, not a hardened multi-user service. Do not expose it through port forwarding, a public reverse proxy or a shared remote host. No trades or account transfers can be performed. Financial decisions should not rely on its unvalidated heuristic scores.
10. Historical replay (research diagnostic)
replay.py re-scores past business days with the current rules and only observations dated on or before each day. Values are the latest selected vintages stored in the database. Later revisions and normal/severe anchors calculated from the complete reference history can therefore influence earlier scores. This is not a point-in-time backtest, even when a forecasting model is trained chronologically.
Reviewed evidence follows its recorded history. The replay, live score, heatmap grid and heatmap detail use the same helper to select the event state known at the requested time. Creation, approval, severity edits, retractions and reapproval take effect at their audit timestamps, and an event cannot appear before its observation or event timestamp. For a legacy event without audits, only the stored state from observed_at onward is available; the app cannot invent missing review history. Historical social samples and exposure histories are not fabricated.
The coverage-limited scenario composite averages whichever A–F scenarios were rated on that date; the live composite requires three rated scenarios. Neither is the headline Stress level. Reference windows help inspect the historical scores and supply anchors for the active-stress component; they are not independent external crisis labels. The script writes CSV, JSON and Markdown reports, and the Reports page displays the latest current replay.
11. Stress level: 100 is the severe boundary
Let S be the active-stress component and V the quality-weighted vulnerability composite. The combined index is:
unscaled_index = 0.30 * V + 0.50 * S + 0.20 * V * S / 100
raw_index = 45 * unscaled_index / min(2008_combined_peak, 2020_combined_peak)
displayed_stress_level = raw_index * 100 / 45Thus raw 45 displays as 100, raw 50 as 111.1, and raw 35 as 77.8. Displayed values can exceed 100; the display does not clamp at the severe boundary. The conversion changes units, not probability. Raw JSON/CSV index units remain compatible, and JSON display_scale metadata identifies the conversion. The expanded input set and all-group aggregation change the calculation and require a new replay. Component, metric, similarity, scenario and scenario-composite scores remain on their original scales. If vulnerability is unavailable, the underlying index falls back to active stress alone and shows that missing component. Without usable active stress, the headline is unavailable.
All S metrics are considered, including monthly and quarterly observations. For eligible non-level-only metrics, the reference anchor uses the smaller risk-oriented peak available in 2008 and 2020 and a normal level from the median outside reference episodes. Design anchors are retained when there is insufficient history or the observed reference peak is not sufficiently severe. Level-only physical, food and health metrics retain their design anchors. The headline S calculation maps observations directly through these level anchors; diagnostic rolling z-scores do not add another boost to it.
Within each active-stress group, anchored scores are averaged using effective quality. Every available group then enters:
intensity = mean(all available active-stress group scores)
breadth = count(group score >= 60) / count(available active-stress groups)
unreferenced_active_stress = 0.70 * intensity + 0.30 * 100 * breadthThere is no top-five selection. Missing groups are omitted from both denominators rather than treated as calm observations. Geographic gaps and changes in available groups can therefore change the interpretation of a score.
This all-group component is normalized by its smaller available reference-window peak and clipped to 0–100. The reference windows are rescored with the same eligible inputs and aggregation rules. The vulnerability/active-stress combination is then calibrated a second time so the smaller combined peak across the 2008 and 2020 windows reaches the severe line. A workspace without reference-window observations explicitly shows an unreferenced fallback score.
Anchors are retrospective and change when the selected history changes, including same-date revisions. Vulnerability counts on its own and also interacts with active stress, but it carries less weight than observed transmission. The display conversion does not alter scenario alerts or evidence requirements.
Descriptive historical resemblance
The heatmap's reference comparisons are separate from the first-crossing timing estimate. Their Bray–Curtis resemblance on shared, observed nonnegative group scores is 100 × (1 − sum(abs(current − reference)) / sum(current + reference)). Missing groups remain unknown; a zero denominator is unavailable. Reference averages use observed values only and require a group on at least half the window's sampled dates, with a minimum of two observations. Each comparison needs at least two shared active-stress groups and two shared vulnerability groups, covering at least half the union of available groups within each component. These are explicit comparability guards, not empirically calibrated warning thresholds. Shared/union coverage is displayed.
Large persistent vulnerability scores can dominate the combined resemblance even when fast stress differs. The detail view therefore shows both component resemblances. No comparison enters warning counts, red/amber streaks, worsening/improving rankings, the Stress Index or its forecast. The colors are neutral blue, and historical cells are unavailable until the named reference window has ended. This removes the misleading impression that an episode's future profile was already a warning before that episode.
Recent commonness fixes today's shared comparison groups, then examines prior weekly dates in the 104-week display that contain all those groups, come after the reference window, and have a contemporaneous raw Stress Index below 45. The current sample is excluded, including on weekends. With at least 12 eligible dates, the detail shows the median, middle-half range and the count at least as similar as today. This is neither a future-outcome success rate nor a failure probability: weekly samples are dependent, weekly sampling can miss intervening severe days, and all calculations use current vintages and retrospective anchors. Similarity is not reduced merely because time has passed without a crisis.
11a. Where we stand: four stages, a verdict and the levels to watch
The Overview opens with a plain-language reading rather than a number. Four stages are scored 0 to 100 from inputs the monitor already carries, ordered by how early each moved before 2008 when the whole record is re-scored with today's rules: the backdrop loads over years (debt to GDP, interest to tax receipts, corporate and household leverage, equity to GDP, term premium, foreign demand, index stretch), credit turns over quarters (lending standards, bank loans to nonbanks, the curve, delinquencies), funding stress appears over months (the Chicago and St. Louis Fed gauges, credit spreads, repo, the sell-America signature), and the break happens in weeks (volatility, risk assets, rates shock). A stage counts as lit at 50. The verdict comes from the configuration of the stages rather than the index level: calm, loaded but not turning, turning, funding stress, breaking.
What the record supports is narrower than the panel might suggest, and the app says so on the Track record page. The *absolute level* of the backdrop stage predicts nothing: it has risen secularly with debt and valuation, has sat above 50 continuously since 2020 without a collapse, and over 2007 to 2026 an absolute-level rule fired on a lower share of months that were actually followed by a severe reading than the unconditional base rate. It sizes a fall; it does not time one. The one trigger with real lift is a rise in the funding stage of twenty points over twelve months. Since 2007 it fired in 18 of 195 tested months, in six runs: two thirds of the months it fired were followed by a severe reading within two years against a base rate near a quarter, a lift of about three times, and it first fired 91 days before the 2008 collapse window opened. It also gave four false alarms, in December 2011, June 2012, late 2015 and early 2016. Two collapses in nineteen years cannot validate anything statistically; these are base rates conditioned on a configuration, reported with their failures.
Episodes carry an origin so the trigger is judged on what it can see. Thirteen built inside the financial system through leverage, lending and funding. Two arrived from outside it, the 2020 pandemic and the April 2025 tariff announcement, and no financial series led either. A pandemic is not a test of a credit-cycle model.
For those outside shocks the app records the only warning they ever gave: the lag between the public signal that the thing was real and the market repricing. The World Health Organization declared a Public Health Emergency of International Concern on 30 January 2020; the S&P 500 then rose for twenty more days to an all-time high on 19 February before falling 34% by 23 March. The declaration was a better exit than the high that followed it. The April 2025 tariff schedule, announced after the close, gave no lag at all. That table sits beside the trigger's history, and it is why the WHO and CDC feeds are in the evidence ledger.
Beside the stages the Overview lists lines to watch in their own units rather than as scores: the 30-year and 10-year yields, the yen, Brent, the VIX, the repo spread over the Fed floor, the Baa spread, lending standards, the St. Louis Fed index, Lake Powell's operating floor and the funding alarm itself. Each shows today's value, the line, the distance and what crossing it would mean. The quiet entries are the falsification checklist for the current reading: they are the argument against a crisis, stated as numbers a reader can check anywhere.
The written brief is four to six sentences assembled from the same state, in the scale the page displays, where 100 is the severe boundary. It states the level and which stages are lit, the verdict, what moved over the past month with real readings in real units, the funding alarm's position and its full record including false alarms, the closest historical analog, and the nearest line to watch. No judgement is introduced that is not already in the numbers, and every claim points at a row the reader can open.
12. First-crossing timing
The exploratory timing estimate on Details asks when the app's Stress level first reaches 100 (raw 45); it also reports a higher comparison threshold, 111.1 (raw 50). It uses complete future paths from up to 40 comparable historical starts, spaced at least 21 business days apart. Every selected one-year outcome must have ended before the forecast date.
Matches use the same eligible vulnerability and active-stress measurement groups, plus index level and recent changes. Financial precursors, climate, water, food and health now enter through V/S groups; promoted inputs are not repeated as separate precursor/climate/health context features. Feature families receive equal weight so a family with many groups does not dominate solely through its size. Missing or stale inputs are omitted, and insufficient comparable history leaves the result unavailable. Matching uses the grouped metric scores; headline active stress additionally applies its reference normalization and breadth formula. News, social and exposure histories are not manufactured to fill gaps.
One set of analog paths supplies both thresholds and all horizons. First crossings are divided into mutually exclusive windows: within 1 month (21 business days), months 2–3 (through 63), months 4–6 (through 126), months 7–12 (through 252), and no crossing within 12 months. The output also gives cumulative shares. A threshold already reached today is identified separately rather than assigned a future onset date.
These are exploratory historical analog shares, not an exact collapse date. The check compares broad-profile matches with index level/trend alone and with the historical base rate on the same comparable issue dates. Brier scores measure threshold-occurrence error; positive skill means lower error against that comparator. More inputs do not automatically improve prediction, and neither comparator performance nor added context proves causality.
Counts of analog starts, nonoverlapping one-year windows and distinct crossing spells accompany the estimates. Starts spaced one month apart still have overlapping one-year outcomes. Severe episodes are sparse; a large row count does not supply a large number of independent crises. Current vintages and retrospectively defined anchors also limit these chronological checks.
The timing payload separately reports readiness, including warning_ready: false, support counts, three-month warning checks and reasons. Availability means the algorithm can compute historical shares, not that a validated warning system exists. Positive Brier skill alone does not establish useful lead time. Unavailable, insufficient and already-reached cases remain distinct. This presentation change does not tune the historical shares to a desired answer.
13. Other research views
Level fans and ridge logistic estimates. The alternative estimators forecast the underlying index 21, 63 and 126 business days ahead. Level fans use historical future levels conditional on the current level band, 20-day trend and, where sufficient history exists, the prior year's vulnerability change. Logistic estimates use index level, changes, breadth, vulnerability, confirmations and a recent peak. Threshold occurrence counts future observations and excludes the starting day. Independent horizon fits are reconciled so longer horizons cannot have lower occurrence estimates and higher thresholds cannot have higher estimates. These answer a different question from the first-crossing analog model.
Chronological checks train on outcomes matured before the test year. They show 10th–90th band coverage, mean absolute error of the predicted median against persistence and climatology, Brier error, reliability and alert diagnostics. The legacy JSON field median_abs_error means the mean absolute error of the median prediction. Headline values and their errors are converted to displayed Stress-level units in presentation; raw output fields remain compatible. No fixed historical skill or peak is assumed in this document: inspect the current replay results.
Two-year projection. The alternative monthly block bootstrap draws changes in active stress and vulnerability together, conditioned on the simulated active-stress band, and recombines the underlying index on each path. A model month is 21 business-day observations. Historical intervals stay fixed when data is appended; the latest partial interval supplies today's starting state but is excluded from training. Future within-month highs enter threshold outcomes; today's starting value does not count as a future hit. Checkpoints cover 6, 12, 18 and 24 months, with chronological diagnostics for matured 12- and 24-month outcomes. Era and El Niño comparisons restrict block starting months; subsequent months in a block may cross a regime boundary. The record contains few independent two-year windows, so these projections are descriptive and sensitive to the chosen history.
Weather, water, agriculture and health. ENSO magnitude, temperature exposure, reservoir elevations, drought area, observed Indian monsoon deficit, international food-price pressure and respiratory surveillance enter V. ONI and weekly Niño 3.4 share one ENSO group. Food and cereal annual changes share one food-price group; the raw monthly indices are upstream inputs, not additional votes. Reservoir readings share a water group, state drought readings share a drought group, and the three health series share a health group. The Treasury yield curve is also a scored vulnerability input; CISA KEV additions enter vulnerability as operational exposure.
ENSO alters the odds of regional rainfall outcomes rather than determining them. Both El Niño and La Niña can be associated with regional droughts or floods; ENSO magnitude alone does not establish Indian drought, harvest loss or financial contagion. IMD supplies the actual national cumulative monsoon departure, which does not establish local crop losses and does not capture localized flooding when the national departure is positive. FAO prices measure international commodity-price pressure, not crop output or household retail inflation. The combination covers selected places and measurements, not a complete global physical-risk system.
All physical and health scores use level-only design anchors. Examples are 0 to 30% annual FAO price inflation and 0 to −30% IMD monsoon departure. Those score anchors and the group weights are heuristics, not validated economic-impact functions. Hydrological thresholds can describe reservoirs without establishing the probability or timing of a market collapse. Grouping related measures reduces duplicate votes but does not eliminate correlation or prove incremental predictive value. El Niño comparisons in the separate bootstrap research use observed ONI history and make no forecast of an unprecedented climate event.
Primary background: NOAA on ENSO and the Indian monsoon, WMO on both ENSO phases, FAO food-price methodology and releases, and IMD observed cumulative rainfall.
Unused research records. Holdings-specific rules, social aggregates and exposure edges are not components of the combined headline. Their unused dashboard views are hidden while stored observations and separate research capabilities are retained. A holdings-related raw price can still appear if it actually supplies an eligible scored factor; usage follows the selected calculation, not the asset's label.
Overview heatmap and date slider. The Overview uses the daily headline replay, with exactly the same conversion to 100 = severe as the current reading. The three-month and two-year views preserve individual daily spikes; the All history view shows the maximum observed daily index in each calendar month and selects its peak date on click. Monthly color is a peak, not an average or the last day. The slider still selects individual dates. Missing weekdays remain explicit and invalidate any five- or twenty-weekday change crossing the gap. A current weekend reading is appended on its own date rather than replacing Friday, and has no weekday trend. Historical selections are labeled current-vintage reconstructions, not forecasts available at that date. Date selections are normalized when a rolling window advances, and background refreshes rerender the slider while preserving its selected date and keyboard focus. The Overview does not add historical resemblance, event counts or individual pathway indices to the headline.
Evolution heatmap on Details. The headline row uses 100 = severe; its red threshold starts at that boundary. Active stress, vulnerability, groups, scenarios, relevant measured-input rows and similarities retain their component scales. Unused holdings rows are hidden. Heatmap dates are rescored from the same selected observations and audited evidence as replay; pathway cells remain blank below the coverage floor. Broad-profile timing is separate from the map's descriptive resemblance to reference windows.
14. September 9, 2026 calculation and refresh update
Automatic updates collect enabled providers, then rebuild the daily replay and forecasting products from the stored observations. Recompute forecasts runs that calculation without provider collection. The app marks saved research stale after observation revisions, evidence audits, relevant settings changes, calculation-version changes or an outdated session, and replaces results only after a completed consistent computation.
The default-enabled physical_enabled collector checks the public FAO and IMD feeds on each hourly update. FAO records use completed calendar-month ends and annual changes compare the same month one year earlier, including leap-year February. Missing comparators are not substituted. IMD records keep the published cumulative June–September period's end date and national aggregate; state percentages are not averaged. Retrieval time is not assigned as an earlier publication time. Failures are recorded independently and preserve prior observations. Outside June–September, IMD is marked inactive with a seasonal explanation and cached readings are retained.
Cache keys now include selected observation values, provenance and disagreement, plus relevant settings and evidence history. Same-date updates no longer rely on the number or date range of rows changing. Review history is preloaded once for replay rather than queried on every date.
Earlier build notes and docs/TEST_REPORT.md describe their own recorded checks; they do not validate the new display scale or timing method. Current regression checks cover revised prices and anchors, cache changes after event/settings edits, audited review/retraction timing, and agreement between live, replay and heatmap scoring. These software checks establish calculation consistency, not predictive reliability.
15. Long history through the latest trading session
The History to today page uses actual daily S&P price observations from Yahoo Finance, including provider predecessor history from December 1927. One consistent price series is used throughout; French total returns are not spliced into it. S&P identifies the modern index launch as March 4, 1957 and describes earlier values as retrospective history. This price series excludes dividends. Historical completed months use the last observed trading-session close, retaining its actual observation date alongside the calendar-month label. FRED supplies Baa and Aaa corporate yields, industrial production, and unadjusted CPI (CPIAUCNS, whose history reaches the Depression). No modern Stress Index values are invented for earlier eras. The archive does not change its severe boundary of 100.
The common model uses trailing twelve-month equity returns, the Baa–Aaa spread, twelve-month industrial-production growth and inflation. Bond yields lag one completed month; production and prices lag two. These are approximate publication lags applied to currently revised data, not a verified historical information set. CAPE remains optional context and is excluded from the common model. Missing measurements remain missing, and incomplete calendar months are excluded from training and validation.
The research outcome is the first future month-end price reading at least 20% below the issue-month price level, within six or twelve calendar months. It differs from a peak-to-trough bear market, an NBER recession and the user's primary Stress Index target. Monthly data can miss a sharp loss that recovers within the month. The single regularized logistic hazard generates a coherent first-crossing path, so cumulative twelve-month loss estimates cannot be below six-month estimates.
Chronological tests freeze a model for each of six predefined eras: 1929–1945, 1946–1969, 1970–1989, 1990–2006, 2007–2019 and 2020 onward. Fitting uses only issue dates whose full twelve-month outcomes end before the test era starts. The model requires 60 matured origins and both outcome classes. With an equity archive beginning in December 1927, the first era cannot be tested; the Depression is observed history and contributes to later training, not a claimed pre-1929 prediction. The latest model uses only outcomes matured before its own issue month. Parameters and feature scaling are fitted inside the training period. Revisions, approximate release lags and structural differences across eras still limit these retrospective checks.
The page compares Brier errors against each fold's training base rate on exactly the same evaluated issue dates. Warning tables show capture at 10%, 25% and 50% alert thresholds, median lead among captured spells, and false/all alert months. All outcomes overlap. First-hit months are grouped into operational loss spells until more than twelve months pass without another hit; this count is not a count of independent economic crises. Advance-warning credit requires the full preceding evaluation window and a forecast before the spell begins; later alerts inside a spell receive no advance-warning credit.
The market-return heatmap shows actual monthly price returns. The latest outlined tile shows the actual dated observation within an unfinished month, with its month-to-date price change. Its alternative Forecasts layer shows held-out twelve-month estimates for historical months and a separately labeled provisional estimate for the current unfinished month, with missing forecasts striped. A monthly slider displays the matched six- and twelve-month outcomes separately from the earlier model estimates. Named reference periods also show the lowest return relative to the preceding month across the whole period; this is explicitly different from peak-to-trough loss or the fixed forecast horizon. Quieter comparison periods remain in the table.
Hourly collection and recomputation run every day while the local app is running and the computer is awake. The scheduler handles only the latest due hour after startup or wake, prevents overlapping runs, persists completed dispatch slots and distinguishes repeated daylight-saving hours by UTC offset. The interface shows automatic update status and has no Refresh Data or Update Archive button. Settings can pause updates or use selected weekday times. Macro archive sources are checked weekly; the S&P source has at most a fifteen-minute successful cache, so hourly updates check it again. Source failures retain prior observations and are reported as partial job results. Raw downloads, hashes and dated metadata remain under history/. Model/source/calendar changes invalidate saved research. Read-only endpoints expose the report and completed-month input CSV; the separate current observation remains present in the report JSON.
The current observation keeps the actual equity-session date and quote timestamp. A quote dated September 9 is never stored as September 30, and it never enters completed-month training or historical tests. The current provisional calculation uses that observed price, an actual same-series session on or just before the prior-year anniversary, the preceding month's credit spread, and production/inflation with two-month lags. Only outcomes completed before the actual issue date enter its fit. Its first future interval may be shorter than a month: the target is the next six or twelve month-end readings, with explicit end dates. Historical month-end tests do not establish validity for this intramonth application. Missing dated predictors or inconsistent series/timestamps make it unavailable rather than fabricated. A weekend or stale quote retains its actual trading-session date.
Primary source documentation: Yahoo Finance S&P price history, S&P index methodology, Baa yields, Aaa yields, industrial production, and unadjusted CPI.
16. Forecast input visibility
The catalog contains 124 metrics, including scored factors, raw inputs to derived factors and records used only by other research. Details and dashboard market views show metrics actually used in the current combined headline, either directly or upstream. The usage record separately identifies unavailable and unused inputs so gaps remain inspectable without displaying unrelated records as drivers. A zero score can still be used; missing is not zero.
Direct membership follows the headline's actual contribution list. Upstream membership requires an observed parent and confirmed computed provenance for a currently included derived factor. An imported derived number does not automatically establish that the app used its usual raw parents. Where alternative sources cannot be resolved from stored provenance, the app does not claim both were selected. Upstream inputs receive no extra contribution points.
All eligible V/S groups also enter the severe-index timing profile; weather, health and the curve do not receive duplicate context dimensions. Historical matches may have fewer available groups. The separate long historical equity-loss model retains exactly four predictors: equity momentum, credit spreads, industrial-production growth and inflation. Expanding the modern headline does not add weather or other modern measurements to the Depression-era model or establish their forecasting value.
Dated heat strips and observed market declines
The Overview and historical archive use a shared chronological axis. Overview colors and its trace encode the same Stress Index, with red beginning at 100. The all-history overview aggregates to monthly maxima and preserves the actual date of each peak. A selected non-peak day does not receive an invented peak-value dot. The archive uses monthly price returns or dated model estimates, never the Stress Index scale. Its unfinished current month retains the actual observation date and provisional status. Missing data is striped and does not connect across a missing monthly interval.
The event catalog records sourced exact dates. Editorial context windows are separate fields and do not identify the duration or onset of a crisis. Monthly observations, historical event dates, forecast issue dates, and later observed losses remain distinct. The century view limits labeled pins for legibility; every cataloged event remains available through the event selector. Markers do not enter forecast fitting or alter stress scores.
The market-drop browser derives a broader list from the cached S&P price record: confirmed close-to-close declines of at least 5%, distinct running-high-to-recovery drawdowns of at least 10%, and completed monthly losses of at least 10%. Daily pairs more than seven calendar days apart are excluded; missing or invalid prices break pair comparisons. Drawdowns carry data-gap and left-censor flags. A further selloff before the old peak is recovered stays within that drawdown episode, so these are not independent economic-crisis counts. Recovery is the first observed close reaching the old peak. Open episodes and retrospective troughs are labeled; gaps may conceal an earlier recovery or deeper low. Provisional quotes are excluded from daily-drop and drawdown calculations, while the archive still shows today's separately dated quote. All measures exclude dividends and preserve the pre-1957 retrospective-index limitation.