Methodology
What our data actually is, how we compute the signals, and which conventions we follow when something is ambiguous.
What this dashboard is
The Distributed is a one-stop financial intelligence dashboard that unifies mainstream market data (live prices, sector heatmaps, earnings calendars, fundamentals) with alternative-data signals the rest of the market usually pays for: insider trades from SEC Form 4 filings, Reddit mention volume and sentiment, coordinated-activity flagging, and cross-ticker correlation analysis.
Each surface in the product answers a specific question:
- Markets (home) — what's notable today across every signal we have, on one page.
- Insider — who's buying or selling their own company's stock, with cluster detection for unusual conviction.
- Earnings — what reports are coming up and what the actuals were, sourced direct from SEC XBRL.
- Screener — combine signals (Reddit volume + insider activity + upcoming earnings) to surface tickers worth a closer look.
- Correlations — which tickers' Reddit mentions move together, plus a search to filter the full correlation set.
- Asset profile pages (
/stocks/AAPL,/etfs/SPY,/commodities/oil,/macro/cpi) — every signal we have for one asset, in one place.
We don't hide methodology. Every metric on the site links back to this page, and every signal explains its source, its window, and its known limitations.
Insider trading
Insider transactions come directly from SEC EDGAR Form 4. We pull the raw filing index, parse the Form 4 XML, and store one row per transaction. There is no third-party data vendor in this path.
The transaction code surfaced in the table is the literal SEC code (e.g. P for open-market purchase, S for sale, A for award). Cluster buys are computed in a Postgres function: cluster_buys_last_7d groups by ticker, requires at least two distinct insider names with buys in the trailing 7 days, and aggregates total purchase value.
Latency: filings appear here within ~10 minutes of EDGAR acceptance. There is a structural ~2-day lag from trade date to filing imposed by Section 16 — that is a regulator window, not a pipeline delay.
Congressional trade performance
The "Estimated trade performance" charts on the Congress board and each politician's profile simulate a hypothetical portfolio that mirrors the member's disclosed trades, benchmarked against a SPY buy-and-hold fed the same cash flows. It is an estimate by construction, because STOCK Act filings do not disclose exact amounts:
- Trade size assumes the midpoint of the disclosed dollar range (a "$1,001 – $15,000" filing is simulated as $8,000.50). The disclosed range is the only real datum; we never present the midpoint itself as a dollar figure anywhere on the site.
- Entry timing uses the reported trade date, snapped forward to the nearest trading session. Filings may lag the trade by up to 45 days, so recent activity can be missing entirely.
- Coverage is stocks and ETFs with a priced ticker only — bonds, options, private funds, exchanges and unpriced disclosures are excluded from the simulation (the trades table still lists them).
- Returns are price returns — dividends are excluded on both the portfolio and the SPY benchmark.
- The SPY mirror receives the same buy amounts on the same days and simply holds. Sells are deliberately not mirrored: the benchmark answers "what if this cash had bought and held SPY instead," not "what if SPY were traded with the member's timing."
Buy and sell markers on a profile chart are the disclosed trades that actually entered the simulation, placed on the session where their effect lands. None of this is investment advice, and it is never a recommendation to copy anyone's trades.
Earnings
Calendar dates are sourced from SEC 8-K filings in which the issuer announces an upcoming earnings call. Where multiple 8-Ks reference the same fiscal period, the most recently filed prevails.
Actuals (EPS, revenue, segment data) are extracted from the company's 10-Q / 10-K XBRL tags once the period closes. We never approximate from press releases — if a number is missing it is because the issuer has not yet filed.
We do not show analyst estimates. Estimate data is licensed by a small number of vendors at high cost, the consensus methodology varies between vendors, and the estimates themselves create a strong anchoring bias when reading actuals. Showing them would force us into a licensing model that contradicts the sources-direct philosophy of the rest of the product. Surprise calculations relative to estimate are out of scope until we can source estimate data with provenance we trust.
Reddit signals
The Reddit data layer reads from a fine-tuned NER pipeline that classifies each post and comment for stock tickers, companies, ETFs, commodities, and macro indicators. The source subreddits are the finance communities we ingest — currently r/wallstreetbets, r/ausfinance, r/stocks, r/investing, r/etfs, r/options, r/asx_bets, r/cryptocurrency, r/asx and r/commodities. The set grows as coverage expands, so treat any count you see quoted elsewhere as of its date.
Mention ratio (the “↗ 3.2x” chip): today's mention count divided by the trailing 30-day baseline mean of daily mention counts for that ticker. We exclude today's value from the baseline so a quiet ticker that just woke up does not deflate its own denominator. Tickers with no mentions in the prior 30 days return no ratio (we will not invent a baseline of 1).
A ratio at or above 1.5x is rendered as trending; at or above 2x we additionally outline the chip in the brand accent to visually separate genuine spikes from drift.
Sentiment uses the FinBERT ONNX INT8 quantized model, scored per mention. Known limitation: WSB sarcasm degrades classifier accuracy by ~20 percentage points relative to formal news copy. We surface the raw label distribution so users can judge whether a flip is a real mood shift or a meme cycle.
Fear & Greed composite
The Fear & Greed composite is a 0-100 weighted mean of seven signals, each percentile-ranked against its own trailing two-year distribution before weighting. Missing inputs are dropped and the remaining weights renormalize — the composite never substitutes a placeholder for a signal that has no data. Five signals are mainstream market data (momentum, volatility, safe-haven flows, credit spreads, the yield curve); two — Reddit sentiment and Reddit mention velocity— come from our own NER pipeline and are the differentiator versus CNN's index.
Safe haven demand
Stocks' trailing 20-session return against Treasuries'. Stocks outrunning bonds reads as Greed (money moving into risk); bonds outrunning stocks reads as Fear (a flight to safety).
Junk bond demand
The high-yield credit spread — the extra yield investors demand to hold risky bonds over safe ones. Narrow spreads signal risk appetite (Greed); widening spreads signal Fear.
Volatility
The VIX against its own trailing trend. Trading below trend (calm) reads as Greed; spikes above trend read as Fear.
Market momentum
The S&P 500 against its 125-day moving average. Above the average is positive momentum (Greed); below it, momentum has turned (Fear).
Yield curve
The 10-year minus 2-year Treasury spread (2s10s), the same series documented under Correlations' sibling yield curve board. A steeper curve leans Greed; a flat or inverted curve leans Fear.
Reddit sentiment
Mention-weighted FinBERT sentiment across the subreddits we track — see Reddit signals above for the model and its known limitations. Net-bullish reads as Greed, net-bearish as Fear.
Reddit mention velocity
Total daily mention volume against its 30-day baseline (the same baseline convention documented under Reddit signals above). A surge in chatter reads as euphoric Greed; a collapse reads as apathetic Fear.
Correlations
All correlation coefficients on this site are Pearson correlation over a 30-day rolling window of daily values, unless the chart caption states otherwise. We use Pearson rather than Spearman because most pairs we plot (price vs mentions, mentions vs sentiment) approximate continuous distributions over 30 samples; rank correlation throws away too much information.
Sample size disclosure: any reported r value with n < 20 daily observations is rendered with a warning glyph and the chart caption explicitly names the sample size. We do not display r values for n < 10 — there is not enough data for the statistic to be meaningful, and showing it would be misleading. Confidence intervals on r are computed via Fisher's z-transform.
Coordinated activity
A ticker is flagged for a given day when it drew at least 30 tracked mentions and the distinct accounts behind them numbered under 15% of the mention count — for example, 200 mentions from fewer than 30 accounts. Both figures come from the daily per-entity rollup of the posts and comments we ingest. A day with no author data is never flagged: no figure, no claim.
Two supporting figures accompany every flag as evidence, not criteria: the day's mentions against that ticker's own trailing 14-day average, and the subreddit that carried the largest share of the day's mentions. A concentration flag on ten times the usual volume reads very differently from one on a chronically quiet name.
On the coordinated-activity board, consecutive flagged days for one ticker — bridging quiet gaps of up to seven days — collapse into a single episode, so a two-week campaign reads as one story rather than thirteen rows.
A flag is a measured pattern, not an accusation. Earnings threads, one viral post, or a single very active commenter can produce the same shape, and intent is not something these figures can state. Account names are never published; everything shown is an aggregate.