US-stock skilled-LLM Agents
ScannerDashboard
GitHub ↗

US-stock skilled-LLM Agents

A suite of 35 LLM agents for US equities — 24 orthogonal signal & fusion agents (fundamentals, events, insider, VIX, congress…) plus 11 placebo controls as a built-in lie-detector. Every backtest is lookahead-free and dual-benchmarked; an honest significance gate (Deflated Sharpe) separates skill from luck. The Scanner turns it into one screen: hold the market by default, tilt only on credible signals. Every LLM call is cost-ledgered; eval is a service, not a script.

Task 1

Browser Automation Agent

Natural-language task → explicit state machine PLAN · LOCATE · ACT · VERIFY · DIAGNOSE. Self-correction via typed root-cause classification, self-maintenance via three-pronged locator (CSS → ARIA role → visible text). Recovery proven by deterministic fault injection.

Try it →

Task 2

SEC 10-K Item Extractor

Paste any EDGAR 10-K URL → layered pipeline: L1 anchor → L2 structural → L3 LLM self-consistency. Each layer fires only when cheaper layers fall short — most filings extract at $0 LLM cost. Platt-calibrated confidence; quarantine on low confidence rather than emitting wrong data.

Try it →

Task 3

Fundamentals → Strategy → Backtest (built on Task 2)

Enter a ticker → its latest 10-K runs through Task 2 → an LLM forms a thesis and picks one executable strategy from a fixed menu, grounded in the filing (with citations). A filing-date-aligned backtest then tests it — signals act only after the 10-K was public, so there is no lookahead — shown as a candlestick with entry/exit markers, an equity curve vs buy-and-hold, and honest metrics (losses included). Quarantined 10-Ks are refused, not strategised.

Try it →

Task 4

Technicals → Strategy → Backtest

Enter a ticker → a snapshot of indicators (RSI, MACD, moving averages, Bollinger, Donchian, volume) computed strictly as-of the most recent close → an LLM picks one executable strategy from a fixed technical menu, grounded in those readings → a lookahead-free backtest over the trailing ~3 years. Signals act on the next bar's open.

Try it →

Task 5

Ensemble — Fundamental + Technical arbitration

Enter a ticker → the fundamental agent (Task 3) and the technical agent (Task 4) run over one common window → an LLM arbiter fuses them, picking one combine policy from a fixed menu (AND / OR / weighted / gated / defer) from each agent's reasoning — not its returns, so the policy isn't fit to the test window. The combined position is backtested vs buy-and-hold and the S&P 500. Degrades to technical-only when there is no usable 10-K.

Try it →

Task 6

Insider (SEC Form 4) → Strategy → Backtest

Enter a ticker → its SEC Form 4 filings are fetched and parsed into open-market insider transactions, keyed off each filing's filing date (not the trade date) so the backtest can only act once public. An LLM picks one strategy from a fixed insider-signal menu (cluster buying / net-$ buying / …) grounded in the as-of readings, then a lookahead-free backtest runs vs buy-and-hold and the S&P 500. Only open-market buys/sales count — grants and option exercises are excluded, and selling is a weak exit signal, not a short.

Try it →

Task 7

Peer / Sector Relative Strength → Strategy → Backtest

Enter a ticker → we resolve its sector ETF from the SEC SIC code (S&P 500 fallback) and compute a relative-strength series (stock ÷ sector), strictly as-of the latest close. An LLM picks one strategy from a fixed RS menu (uptrend / breakout / momentum); a lookahead-free backtest then holds the stock long/flat with RS deciding only when to be long — vs buy-and-hold and the S&P 500. Surfaces nuances like “beating the market but lagging its own sector.”

Try it →

Task 10

Portfolio / Risk Sizing — the capstone

Enter a watchlist → each name's long/flat signal comes from the Task 4 agent; an LLM picks one sizing policy (equal-weight / inverse-vol / risk-parity / signal-proportional + single-name cap, gross cap, vol target, rebalance) from as-of universe stats — not per-name weights, which are deterministic. A lookahead-free portfolio backtest then allocates across names vs an equal-weight basket and the S&P 500. The piece that turns the signal agents into a sized, risk-controlled book.

Try it →

Task 8

Earnings (SEC 8-K) → Strategy → PEAD Backtest

Enter a ticker → its recent earnings press releases (8-K Item 2.02 / Ex-99.1) are fetched; an LLM classifies each as-of its filing date (sentiment / guidance / beat-miss, with citations); then it trades the post-earnings drift — lookahead-free, acting only on the open after each filing, vs buy-and-hold and the S&P 500. Reads the press release, not the live Q&A transcript (source is pluggable).

Try it →

Task 9

Institutional (13F) superinvestor tracking → Strategy → Backtest

Enter a ticker → we track a curated set of well-known managers (Berkshire, Baupost, Pershing Square, …) via their SEC 13F-HR filings and follow whether they're accumulating the name. An LLM picks a follow-the- smart-money strategy; the backtest is keyed off the 13F filing date (~45-day lag → a lookahead-safe but slow confirmation signal). Curated funds only, matched by issuer name — both surfaced honestly.

Try it →

Task 11

Fundamentals Trend (XBRL)

Structured quarterly financials from SEC XBRL → lookahead-safe fundamental momentum (YoY revenue/ earnings growth + margin trend, keyed off the filing date, as-originally-reported) → backtest. Task 3 reads the 10-K text; this reads the numbers.

Try it →

Task 12

Seasonality / Calendar Effects

Month-of-year returns, sell-in-May, turn-of-month → an LLM picks a calendar rule (lookahead-free to execute). Honest that the pattern is in-sample — weak signals default to buy-and-hold rather than overfit.

Try it →

Task 13

Overnight vs Intraday (Gap)

Splits returns into the overnight (close→open) vs intraday (open→close) move — the documented anomaly that most US-equity return accrues overnight — then backtests a participation rule honestly net of the daily round-trip cost that usually erases the gross edge.

Try it →

Task 14

Volatility Regime / Risk Mgmt

Trailing realized volatility + percentile → a vol-managed long/flat rule: participate when calm, step aside when vol spikes. Judged on risk-adjusted terms.

Try it →

Task 15

Share Buybacks (XBRL)

Falling diluted share count from SEC XBRL = net buybacks → follow sustained repurchases. Lookahead-safe, quarterly.

Try it →

Task 16

Short Pressure / Squeeze (FINRA)

FINRA daily short-volume ratio (weekly-sampled, cached) → a squeeze / low-short rule. Honestly flagged: this is short volume (incl. market-maker hedging), not short interest — free historical short-interest doesn't exist, so it's the closest free proxy.

Try it →

Task 17

Fundamental Quality (XBRL)

Three classic free quality factors from SEC XBRL — Piotroski F-Score, Sloan accruals(earnings quality), and the asset-growth anomaly — point-in-time, filing-date keyed; the LLM picks the factor or a composite. Numbers, where T3 reads the 10-K text.

Try it →

Task 18

Corporate Events (8-K / 13D)

Schedule 13D activist stakes (positive drift) + red flags (dilution, late filings, auditor changes, delisting, adverse 8-K 5.02 exec departures the LLM reads from the text). Keyed off the filing date; ride activist drift, stand aside on red flags.

Try it →

Task 19

Price Anomalies

Three documented price anomalies (prices only): 52-week-high momentum, MAX/lotteryavoidance, and tax-loss reversal (Jan effect). LLM picks one; trailing-window + calendar → lookahead-free.

Try it →

Task 20

VIX Regime Gate

The CBOE VIX term structure (^VIX vs ^VIX3M) as a regime switch: hold long in contango (calm), step aside to cash when the curve inverts or VIX spikes. Same-day signal → next-open fill, lookahead-free.

Try it →

Task 21

Cross-sectional Factor Ranker

Give it a watchlist; the LLM picks one long-only cross-sectional factor (12-1 momentum, low-vol, near-52w-high, or short-term reversal). Each rebalance the universe is ranked on trailing data and the top-N is held — benchmarked against the equal-weight basket of all names and the S&P 500.

Try it →

Task 22

Congressional Trading

US lawmakers' disclosed trades (STOCK Act), keyed to the disclosure date so the edge is post-disclosure drift. Pluggable data: free House-PTR parse, or a Quiver/FMP key for full coverage.

Try it →

Task 23

Pairs Trading (stat-arb)

Two correlated names → a spread z-score → a market-neutral mean-reversion bet (long the cheap leg, short the rich). The suite's one long-short strategy; judge it on Sharpe, not raw return.

Try it →

Task 24

Earnings Contagion

A bellwether's earnings move its peers before they report. Classify the bellwether's 8-Ks (reusing Task 8) and trade the peer in the short read-across window — keyed to the bellwether's filing date, lookahead-free.

Try it →

Task 25

Financial Astrology · Control

Placebo arm. Mercury retrograde / moon phase / planetary aspects through the same honest backtest — to measure the suite's false-positive rate. If it shows alpha, the framework leaks.

Try it →

Task 26

梅花易數 I Ching · Control

Placebo arm. A hexagram cast from the date drives a 體用生剋 rule — and N random seeds give a null distribution the real agents are scored against (a poor-man's Reality Check).

Try it →

Task 27

八字 Four Pillars · Control

Placebo arm. Casts the company's natal chart from its listing date, reads the 日主/喜用神, and holds when the current 流年五行 is favourable — the same honest backtest on a worthless signal.

Try it →

Task 28

紫微斗數 四化飛星 · Control

Placebo arm. Casts the company's full 紫微 命盤 (12 palaces, 14 stars) from its listing date and trades the 四化飛星— does the year's 化祿/化權/化忌 fly into 命宮/財帛/官祿? — through the same honest, lookahead-free backtest.

Try it →

Task 29

四柱推命 (日) · Control

Placebo arm. The Japanese (京都泰山流) reading of the four pillars — driven by 十二運星 and 天中殺 (空亡) (細木数子's axes), not 五行. Same honest backtest, worthless signal.

Try it →

Task 30

七政四餘 · Control

Placebo arm. Real Chinese horoscopic astrology — 七政 (日月+五星) and 四餘 (羅睺/計都/月孛/紫炁) from ephem; signal rides benefic/malefic transits vs the natal Sun. Worthless by design.

Try it →

Task 31

鐵板神數 · Control

Placebo arm (double). The legendary 鐵板神數 條文 book is proprietary, so a deterministic 太玄數 起例 over the natal 四柱 stands in — 命數 → 流年條文 → 吉凶 drives hold/flat. Worthless by design, twice over.

Try it →

Task 32

奇門遁甲 · Control

Placebo arm. One of the 三式 — 八門九宮 起局; hold on a 三吉門, flat on a 凶門. Worthless by design.

Try it →

Task 33

大六壬 · Control

Placebo arm. One of the 三式 — 月將加時 → 用神; hold when it 生扶 the 日主, flat when it 剋洩. Worthless by design.

Try it →

Task 34

太乙神數 · Control

Placebo arm. The third 式 — 太乙積年 → 主算 vs 客算; hold when 主勝, flat when 客勝. Worthless by design.

Try it →

Task 35

Jyotiṣa (Vedic) · Control

Placebo arm. Sidereal Vedic astrology — 9 grahas (tropical − Lahiri ayanāṃśa) + the Vimśottarī Mahādaśā (the exact 120-yr period cycle); hold in a benefic daśā. Worthless by design.

Try it →

What the dashboard shows

The /dashboard surfaces eval pass-rates, recovery-rate, p50/p95 cost & latency, total ledger spend, cost-by-purpose / cost-by-model breakdowns, recent jobs, and the supported / unsupported capability matrix — every number is queried directly from the cost ledger or eval report.json, not estimated.

Full design in PLAN.md · ADRs in docs/adr · Per-task analysis in docs/analysis.