Methodology
How WhatsGoingOn decides what counts as a thesis, and what counts as being right.
What we track
Reddit. DD and YOLO posts from r/wallstreetbets that clear 20 upvotes, evaluated 20 minutes after posting to let scores settle. Each post is parsed by an LLM (Claude Haiku 4.5, prompt v4) into a ticker, a direction, and a summary of the argument. Tickers are validated against real US-listed symbols. Extractions scoring below 0.60 confidence are hidden by default.
A thesis requires a ticker and a direction. Memes, news links, and commentary without a view produce nothing. One post can produce multiple theses — an argument for AMD and against INTC is two records.
SEC Form 4. Open-market trades by corporate insiders — officers and directors, not funds or trusts filing as 10% owners — parsed directly from EDGAR with no AI involved. Only open-market purchases and sales count (no option exercises, grants, or gifts), the aggregate trade must be at least $250,000, and sales made under a pre-scheduled Rule 10b5-1 plan are excluded entirely. Purchases are recorded bullish, sales bearish, scored on the same scale.
Schedule 13D. When an investor — often an activist fund — crosses 5% ownership of a company with intent to influence it, they must disclose the stake on a Schedule 13D. We record each original 13D as a bullish call by the filer, anchored to the filing time. Amendments (stake increases or exits) and passive 13G filings are not tracked. A 13D discloses shares and percent of class but no purchase price, so these cards show the stake rather than a dollar amount.
How returns are measured
We track the underlying stock, not your P&L. If someone posts a thesis and buys calls, we measure what the stock did — not what the calls did. We're scoring the directional call, not the options structure.
All returns use split- and dividend-adjusted closes, measured from the post timestamp forward.
Seven windows, always shown together: 1 day · 3 days · 1 week · 1 month · 3 months · 6 months · 1 year.
The 1-day window uses the next trading day's close versus the post-day close. Unresolved windows show “—” — we never display projected returns.
There is no composite score. A 70% win rate at one day and 40% at one year describes a specific kind of trader, and a single blended number would erase that.
What counts as a win
| Sentiment | Return | Result |
|---|---|---|
| Bullish | Positive | Win |
| Bullish | Negative | Loss |
| Bearish | Negative | Win |
| Bearish | Positive | Loss |
| Neutral | Any | Excluded |
No magnitude threshold. +0.1% is a win exactly like +40%. Any cutoff is arbitrary and “why 3% and not 2%” is an argument that produces no information. The raw return is displayed next to every win, so you can apply your own bar.
Paper trading
Signed-in users can “tail” a thesis: open a simulated long or short position in the underlying stock. Entries fill at the most recent quote we hold (refreshed roughly every minute during US market hours; last close otherwise), and every fill is timestamped with the quote's actual age. Positions settle at the chosen duration's end using the same adjusted-close-to-adjusted-close math as thesis returns, so splits and dividends don't distort results. Open positions are marked to market against the latest quote without adjustment. Positions in stocks that stop trading are closed at the last available price and flagged. All paper trading is percent-return only — no dollar amounts — and simulated trades never touch real markets.
Paper trading is a simulation. Fills happen at the last observed quote — which may be minutes or hours stale — and ignore spreads, slippage, liquidity, borrow availability, and fees. Short positions assume free, unlimited borrow. Results are hypothetical, would not have been fully achievable with real orders, and are not investment advice.
Insider filings: read as transactions, not predictions
Most insider sales are not directional views. They're tax withholding on vesting equity, or diversification by someone dangerously concentrated in one stock — and the most mechanical kind, scheduled 10b5-1 plan sales, never enter the ledger at all. A CEO selling before a rally was not necessarily wrong; they may not have been expressing a view.
We score the rest consistently because that's the only way to find out whether any of it carries signal. Insider purchases are the more informative side; nobody buys their own stock by accident.
The answer engine
The “what's going on” box on the dashboard synthesizes an answer from the same ledger everything else on this site reads. It uses a language model in exactly two places, and neither of them is allowed to produce a number.
What the model does and does not do. Every figure in an answer — every count, dollar amount, percentage, and win rate — is computed by database queries before the model is involved. The model then does two jobs: classify your question, and judge which of the pre-computed events matter most, writing two or three sentences about them. Before that summary is shown, every number in it is checked against the set of figures the model was actually handed; anything unverifiable — a computed total, an estimated count, a company name recalled from memory — causes the whole sentence set to be discarded and replaced with a plain template built directly from the data. The same check rejects any forward-looking language: the engine describes what has been said and what has happened, never what will. In production this check has already caught the model miscounting its own list. When it fires, you still get a complete answer — just one written by a template instead of a model. The “how this was assembled” expander on every answer says which one you got.
Source weights. Not every artifact starts equal. Each type carries a base weight — an activist stake disclosure 0.90, an insider purchase 0.75, an insider sale 0.45, a Reddit DD post 0.35, a YOLO post 0.20 — reflecting our judgment of how much a fresh one is worth before anything else is known about it. Purchases outrank sales for the reason stated above: nobody buys their own stock by accident. These weights are editorial priors, not measurements. Once each type has 50 or more resolved windows, we will recalibrate them against realized win rates and update this page.
Recency decay. Every artifact's weight halves on a clock suited to its kind: a YOLO post every 3 days, a DD post every 5, a Form 4 every 30, an activist stake every 90. A stake filed three weeks ago still outweighs yesterday's meme — and an activist position that has aged out of the “what changed” list moves to “where things stand” rather than disappearing, because it is still true.
Credibility weighting. When an answer says the crowd is “38% bullish weighted by track record,” each author's vote is scaled by their shrunk 3-month win rate: (wins + 5) / (resolved + 10). The constant pulls small samples toward a coin flip — an author who is 2-for-2 counts as 58%, not 100%. Credibility then adjusts an artifact's weight within a bounded range (0.6× to 1.4×): a bad record discounts an author, it never silences one, and a good record earns emphasis, never a megaphone.
Scoreable versus context. Only artifacts that can be scored against what the stock did — posts, insider trades, activist stakes — enter win rates, weighted sentiment, or either side of a contradiction. Everything else, price moves included, is context: shown, labeled, and permanently excluded from the scoring math. When we add sources that cannot be scored (institutional 13F positions are 45 days stale at publication), they will live on the context side of this wall.
Contradictions. An answer leads with “sources disagree” only when at least two independent providers (Reddit and SEC filings — two filing types from EDGAR don't count) point in opposite directions on the same ticker, both sides carry real weighted mass (at least 0.50), and neither side is more than 5× the other. The model may decline the headline when the arithmetic is technically met but the story isn't there; it can never manufacture one the thresholds didn't fire. These thresholds are provisional and will be re-fit once enough real answers have accumulated to tune them against.
Limitations
Sample size. Many authors have fewer than five theses. A 100% win rate on two resolved windows means nothing. Thesis counts are shown next to every rate — discount accordingly below ~20.
Survivorship bias. Historical posts come from Reddit's own listings, which favor posts the community liked. Deleted posts and accounts are gone. The sample skews toward confident writing, which is not the same as correct writing.
No benchmark. Returns are absolute. A bullish call during a 20% market rally scores as a win regardless of whether it beat an index fund. Benchmark-relative returns are planned, not shipped.
Extraction error. A model reading sarcasm, slang, and deliberate hedging gets some of them wrong. Every thesis links to the original post — if our summary disagrees with the source, trust the source.
Timing lag. Form 4 filings can post days after the transaction date. Reddit posts are anchored to post time, but the first usable price is the next trading-day close.
Informational only. Past performance does not predict future results. Nothing here is investment advice.
Last updated August 2026 · Extraction prompt v4 · Found an error? Tell us — corrections improve the ledger.