Pipeline, from a new token to a label
- Ingest. A WebSocket subscription to the pump.fun program on Solana delivers every create, trade and migration event as it happens. No third-party API sits in between.
- Early score at about 45 seconds from what exists at creation: deployer history, block-0 wallets, metadata, credibility and the local analyst.
- Main score at 5 minutes. The feature window closes, 123 signals are computed on the trades of those five minutes with wallets clustered, and the model gives the probability of rug.
- Automatic label at 1 hour from what really happened after the window (rules below). The token becomes a training example without anyone touching it.
- Reviews at 24 hours and 7 days with full on-chain history: collapses that recovered, insiders that sold late.
- Human votes on any token from Train AI. A human mark always prevails over the automatic one and goes through the AI review queue, which infers the reason and explains it.
- Retrain on the labelled set. A challenger replaces the champion only if it does no worse on the last 24 hours of live tokens.
Database today: 25,000 tokens scored, 23,000 labelled, 2.5M trades, 187,000 wallets with cross-token memory, 179 human marks.
Signals · 123 features
All computed from on-chain data of the first five minutes plus metadata, never from the token's own description taken at face value. Contract signals act as filters; distribution and behaviour carry the weight. Concentration is measured after clustering wallets by creation slot, common funder and co-buying, which is what bundlers try to hide.
| Group | Signals | What they measure |
|---|---|---|
| Contract | 8 | Mint and freeze authority, Token-2022 and dangerous extensions, non-standard supply, Mayhem Mode, USDC quote |
| Deployer | 19 | Initial buy and its size, dev sells inside the window, previous tokens found on-chain and how many died or migrated, wallet age and activity, creation rate |
| Bundle and snipers | 10 | Buys in the creation slot, buys in the first 5 seconds, supply they took and how much they still hold at window close |
| Concentration | 18 | Holders, top-1 and top-10, HHI and Gini, clustered top-10 and the gap versus raw, supply still held by anyone, wallets seen in earlier bundles, snipes or labelled rugs, wallets sharing a funder |
| Trading | 28 | Buys, sells, unique buyers and sellers without dust trades, volume, price change and drawdown, bonding-curve progress, time to first sell, wash trades, holding times, dust trades |
| Context | 9 | Hour and weekday, social links present, name and description length |
| Credibility | 19 | X account, Telegram, website and domain age, ticker colliding with a Jupiter-verified token, reused image, description similar to recent rugs, celebrity keywords |
| Local analyst | 12 | Qwen 2.5 14B reads name, description, links and website text as untrusted data: credibility 0–100, effort, narrative, flags for impersonation, promises, generated text, incoherence, urgency, empty metadata and prompt injection |
Training parity. Signals that only exist in the present (X, Telegram and website checks made today) are excluded from training so that history and live tokens share the same features. Only tokens with at least 3 real buyers at 5 minutes are used for training and evaluation: the ones someone might actually buy.
Dust. Trades below 0.001 SOL are bots probing the curve. They no longer count as buyers, sellers or activity. They are about 13 % of all trades.
Labels · 1 h, 24 h, 7 d
Window: first 300 seconds. Horizon: 3,600 seconds. The reference price is the last price at window close. Every rule is quantitative and uses only trades of the pump.fun program.
| Subtype | Family | Rule | When |
|---|---|---|---|
| honeypot | insiders | Mint or freeze authority active, or dangerous Token-2022 extensions | window close |
| honeypot_suspect | insiders | 15 or more distinct buyers and not one successful sell in the whole hour | 1 h |
| dev_dump | insiders | Price falls 60 % or more and the creator wallet sold half or more of a position of at least 0.5 % of supply | 1 h |
| bundle_dump | insiders | Price falls 60 % or more and the creation-slot wallets, holding 1 % or more, sold half or more | 1 h |
| sniper_dump | insiders | Price falls 60 % or more and the first-5-second buyers, holding 1 % or more, sold half or more | 1 h |
| whale_dump | insiders | Price falls 60 % or more and one or two wallets with 5 % or more of supply made half of all sells | 1 h |
| insider_exit | insiders | The dev (or the bundle) sold 90 % or more of its position and at 1 h holders keep under 2 % of supply. No price drop required: the token was emptied on a thin curve | 1 h |
| price_collapse | collapse | Price falls 60 % or more and none of the above explains who sold | 1 h |
| dead_idle | abandonment | No real trades after the window, or inactive for more than 80 % of the horizon, without a collapse. Counts as rug | 1 h |
| survivor | ok | Neither collapse nor inactivity nor insider exit during the horizon. Provisional until the reviews | 1 h |
| graduated_unknown | excluded | Migrated to PumpSwap without a prior collapse. Post-migration trades are not decoded yet, so no outcome | 1 h |
| recovered | ok | Fell 60 % or more but came back to 80 % of the reference (or above the peak) without an insider dump | 24 h / 7 d |
| insider_dump_recovered | insiders | Insiders dumped and the price still recovered. Rug for the model even if the token lives | 24 h / 7 d |
| slow_rug | insiders | Survived the first hour; later insiders or the 5 largest holders sold half or more and the price ended at 40 % of reference or less | 24 h / 7 d |
| late_collapse | collapse | Survived the first hour and at 7 days sits at 40 % of reference or less without insider sells | 7 d |
Binary label. Insider families and collapses are rug. Abandonment counts as rug (a token nobody can sell into is a loss). Graduated tokens without decoded outcome are excluded from training.
Human marks. A vote from Train AI or the internal dashboard is stored as the human label and never overwritten by the machine. A mark given before the token is 7 days old is provisional: the 7-day review compares it with the real outcome and, if they contradict, returns it to a review queue for a person to look again. The AI review queue infers the reason family and writes an explanation for every human mark.
Risk bands on the site: HIGH ≥ 70 % MEDIUM 35–69 % LOW < 35 %. There is never a green "SAFE" badge. The model's own decision threshold is set by MCC at each training, 0.35 today.
Latent risk: who can sink the price now
Balances are rebuilt from every trade and the bonding-curve reserves are the last known. For each entity, the drop is what the curve (x·y = k) would do if it sold everything right now: the largest holder, the largest cluster, the deployer, the block-0 bundle, the snipers, and the top 10 together. The card shows the worst of the first five. For migrated or USDC-quoted tokens the reserves are estimated and the card says so.
Model and metrics
An ensemble in logit space of three members: an MLP (two hidden layers, 64 and 32, PyTorch on the GPU), a logistic regression with editable coefficients, and XGBoost (300 trees, depth 4). Class-balanced training, early stopping by AUC-PR, decision threshold chosen by MCC, per-feature importance by permutation. A second, smaller model scores at 45 seconds from creation-time features only.
Validation is temporal. The last 20 % of tokens by creation time is held out: train on the past, evaluate on what came after. Hard rules from the reference document (mint authority, bundle above 30 %, clustered top-10 above 70 %, serial rugger…) are kept only if they do not hurt validation AUC-PR.
| Metric · model of 29 Sep 2026 | Value |
|---|---|
| Tokens in train + validation (real sources, ≥ 3 buyers) | — |
| Base rate of rug in validation | — |
| AUC-PR, temporal validation | — |
| ROC-AUC, temporal validation | — |
| MCC, temporal validation | — |
| AUC-PR on live tokens, last 24 h (n = —) | — |
How the live panel is computed. For every live token of the last 24 hours with a label, the probability given at 5 minutes is compared with the label that arrived at 1 hour (or the human mark if there is one). Hits and misses use the model's threshold. It is the same number lenders will be shown before they deposit.
Champion versus challenger. Each retrain produces a challenger evaluated on the same live tokens as the current champion. It is promoted only if its live AUC-PR is not worse by more than 0.005. The reference academic model (XGBoost on 6.4 million tokens, Catching the Rug, 2026) reports AUC-PR 0.80 and MCC 0.39 on a comparable task.
Votes and comments
Today. One vote per wallet per token: Rug, Legit or Unsure. A vote is stored with the wallet that signed in and can be changed. The simple majority of Rug versus Legit sets the token's human label with author "consensus" and sends it to the AI review queue. The creator wallet cannot vote on its own token. Comments are public, tied to the wallet, 500 characters, one every 10 seconds.
planned With the token: consensus at 15 weighted votes and 70 % agreement, vote weight = reputation × √(tokens staked), reputation 0–100 from the last 200 closed votes (everyone starts at 50, a miss subtracts double), levels Rookie, Scout, Analyst and Oracle with reward multipliers ×1 to ×3, and payment only for votes that match the final label. Bait tokens with a known label mixed into the queue. Wallets linked to a creator excluded from voting on its tokens.
Wallet sign-in
No password and no transaction. The site asks the server for a nonce, the wallet signs a message containing it, and the server verifies the ed25519 signature against the wallet's public key. A session token valid for 30 days is issued and kept in your browser. Phantom, Solflare, Backpack and any wallet that speaks the Wallet Standard work; on mobile the site opens inside the wallet's app. The session is what signs votes, comments and the leverage waitlist.
API
Public today, same origin as this site, no keys and no quotas yet. Rate limits, keys per plan and the GET /api/risk/{mint} shortcut for bots and terminals are planned.
GET /api/public/summary counters, model metrics, live 24 h, waitlist (cached 20 s)
GET /api/tokens/{mint} token, features, label, human mark, trades series,
explanation (top contributions), latent risk, predictions
POST /api/analyze {query} analyse a link or mint on demand → {id}
GET /api/analyze/{id} stage, seconds, result | error
GET /api/metrics/live?hours=24 predictions at 5 min vs labels at 1 h, by model version
GET /api/features the signal list with group and description
GET /api/validation/next a resolved token without a human mark
GET /api/tokens/{mint}/votes tally and your vote (with session)
POST /api/tokens/{mint}/votes {vote} rug | legit | unsure (session required)
GET /api/tokens/{mint}/comments newest first
POST /api/tokens/{mint}/comments {text}
GET /api/auth/nonce?wallet= message to sign
POST /api/auth/verify {wallet, message, signature} → session token
WS /ws hello (recent + live), token_new, token_early,
token_featured, token_labeled, status
Broker parameters planned · phases 3 to 5
From the broker spec. None of this is deployed; the numbers are the starting values the simulator and the first weeks on mainnet will calibrate.
pool rate = 8 % + 30 % × utilisation (up to 80 % use, steeper above) loan ≤ (leverage − 1) × collateral loan cap = 2–5 % of token liquidity, shrinking with rug probability delist when rug probability > 40 % auto-close creator wallet sells > 1.5 % supply / 10 min, or cluster > 3 % interest split 70 % lenders · 25 % insurance fund · 5 % treasury insurance target 10 % of live loans; below 4 % every market reduce-only tier D ≤ 10 % of total loans liquidation partial, liquidator earns 1 % withdrawals 24 h cooldown
Four mandatory simulations before mainnet: a rug, a 30 % crash, 200 liquidations, a compromised oracle, all with bad debt under 2 % of the token cap. Prices signed by 3 independent feeders with TWAP, median and deviation checks. No leverage on mainnet until the program has 90 % test coverage.
Stack
| Layer | Today | Planned |
|---|---|---|
| Chain data | Public Solana WebSocket (logsSubscribe on the pump.fun program) for live events; Helius for full histories, on-demand analysis and deployer history; PublicNode for heavy RPC | Own Helius plan with webhooks and funder tracing; Triton or QuickNode as fallback |
| Model | PyTorch (CUDA) + XGBoost ensemble on one consumer GPU, RTX 5080 | Same, on an always-on GPU box |
| Analyst | Qwen 2.5 14B served by Ollama on the same GPU; answers in Spanish and English; reads websites | — |
| Database | SQLite in WAL mode: tokens, trades, features, labels, wallets, predictions, votes, comments, waitlist | PostgreSQL 16 and Redis for scores and the feed |
| API and web | FastAPI serving the JSON API, the WebSocket feed, the internal dashboard and this site (one HTML file, no build step) | Fastify API with keys and quotas; Next.js site; Telegram bot |
| Sign-in | Wallet signature verified with PyNaCl, HMAC session | Solana Wallet Adapter in the Next.js site |
| On-chain programs | None. Nothing is custodied | Anchor: rewards and stake; then markets, positions, pool, insurance fund, oracle, shield |
| Keepers | In-process jobs: recheck at 24 h and 7 d, AI review queue, enrichment, backfill | oracle-feeder, liquidator, order-executor, tier-manager, shield-watch, rug-scorer |
Known limitations today
- Trades after migration to PumpSwap are not decoded: graduated tokens, about 1 % of launches and the most interesting ones, get no outcome and are excluded from training.
- Funder tracing per wallet is off: a bundle funded from an exchange is not grouped, and a dev who moves tokens to another wallet before selling reads as a whale.
- Honeypot detection checks the mint only. There is no sell simulation.
- The live feed depends on the public Solana WebSocket, which drops the connection about once a minute; the source reconnects and deduplicates, but a burst can be missed.
- No API keys, quotas or rate limits. The analysis endpoint spends Helius credits on every call.
- Reputation, levels and rewards are not implemented: they open with the token.
FAQ
Is a low probability a guarantee?
No. It is a probability from a model that is wrong a measurable share of the time. That share is on every page, computed on the last 24 hours of real tokens.
Why did the model call a rug "legit" once?
It happened. A dev sold everything at minute two on a nearly empty curve; the price only fell 11 %, so the 60 % rule labelled it a survivor and the model learned from cases like it. That is why the insider_exit label and the dust filter exist now. Every miss is visible in Live, and every label can be corrected with a vote.
Does Racksx bet against me?
No. Nothing is custodied today, and when leverage arrives the PnL comes from the real market and the loan from other users. Racksx earns a fee and 5 % of interest, never your loss.
When can a lender lose money?
Only when sold collateral does not cover a loan and the insurance fund is empty. Caps, early liquidation, auto-close and the fund exist to make that rare and small, and the history of every covered day will be public.
Why are votes not on-chain?
A vote signed by the wallet and stored in the database costs zero gas and can be corrected. Only stake and payouts will touch the chain, plus a weekly hash of the vote dataset so anyone can check it was not altered.
Does any data leave the server?
Only the chain reads and the calls to Helius. The model and the analyst run on a local GPU.
Is the token a share of the pool or of fees?
No. It is a memecoin with access utility and a bounty for labelling work. Rewards are payment for labelling, not dividends. Holding buys access and discounts, never a share of the pool or of protocol fees.
What about regulation of leveraged products for EU retail?
Legal consultation on MiCA and CNMV starts in phase 3, before the first real USDC in the pool. Terms, risk policy and country geo-blocking are ready before then.