Pre-launch· Token not live yet — no rewards or treasury until launchWhat this means
Methodology

How TRAIN works, and what it can't do

Plain answers to: what TRAIN looks at, what it doesn't know, how comparisons work, how versions are judged, how the future is kept out of the past, and how the token pays for it.

What TRAIN analyzes

TRAIN studies Solana token launches on pump.fun. For each launch it records a snapshot at fixed ages (1m, 5m, 15m, 30m, 1h) made of 36 numbers: the price path, volume and its acceleration, buy versus sell flow, how many wallets traded and how many still hold (estimated from trades), how concentrated those holdings are, what the creator wallet bought and sold, how many tokens that creator launched before, and whether socials were set at launch.

Launches are sampled every minute by a hash of their address, so the dataset includes the many launches that go nowhere, not just the famous ones. So far TRAIN has tracked 3,652 launches and 308,646 trades.

What TRAIN does not know

It does not know the future, who is behind a wallet, private group chats, off-chain coordination, news, or a project's intentions. It cannot tell you whether to buy or sell, and it never will.

Holder concentration in the model is estimated from trade flow (tokens bought minus sold per wallet). It misses transfers between wallets and can't link wallets controlled by one person. Wallet labels such as “insider” or “smart money” are not used, because the evidence for them would be weak. When TRAIN mentions wallets with a good record, it means a positive realized result across launches in its own sample, and says so.

It learns from the first hour of launches. For older or migrated tokens the analysis says its estimates are less reliable.

Facts, estimates and history

Fact
Observed now: trades, prices, balances, creator activity.
Model estimate
A probability from the current TRAIN version.
Historical statistic
What happened to similar launches in the dataset.

Every statement on the Analyze and Talk pages carries one of these labels.

How historical comparables work

To compare a token, TRAIN takes launches in its dataset at the closest snapshot age, puts the main setup measures on the same scale, and picks the nearest ones (about 15% of the pool, at least 10). It then reports what share of them migrated, doubled, fell more than 70%, or were still trading. These are frequencies in a sample. A setup that ended badly 60% of the time still ended well 40% of the time.

What the model is asked

Collapse risk · 1h
Chance the price falls more than 70% below its current level within the next hour.
Survival · 1h
Chance the token is still being traded at all (any trades in the last 15 minutes) one hour from now.
Reaches 2× · 1h
Chance the price at some point doubles from here within the next hour.
1h trajectory
Where the price is one hour from now: down (≤0.6×), flat, or up (≥1.3×).
Creator exit · 1h
Chance the creator wallet sells at least half of what it holds within the next hour.
Collapse risk · 6h
Chance of a >70% fall within six hours.
Survival · 6h
Chance the token is still being traded (any trades in the final hour) six hours from now.
6h trajectory
Down / flat / up six hours from now.
Survival · 24h
Chance the token is still being traded 24 hours from now.
Migration · 24h
Chance the bonding curve completes and the token migrates to PumpSwap within 24 hours.

Definitions: a fall of more than 70% means the lowest traded price within the horizon was at most 30% of the snapshot price. Survival means at least $10 of trading in the final 15 minutes of a 1-hour horizon ($10 in the final hour for 6h/24h). Trajectory: up if the price ends ≥1.3×, down if ≤0.6×. Creator exit: the creator sells at least 50% of what it held at the snapshot. Migration: the bonding curve completes.

Keeping the future out of the past

A model that sees the future during training looks brilliant in testing and is useless in practice. Every snapshot has a hard time boundary. Features may only use trades at or before it, price candles that had closed by it, and creator launches that happened before the token launched. The code that builds features throws away anything later itself, and the database refuses to store a snapshot whose latest input is after its boundary.

What happened next — the price path, migration, creator exits — is stored separately as labels and only uses data strictly after the boundary. A label stays empty until its whole horizon has been observed. Automated tests feed the feature builder future data and check that the output doesn't change.

Live analysis uses the same feature code as training, so the model sees the same kind of inputs in both places.

How versions are judged

Launches are split once, by a hash of their address: 70% training, 10% validation, 20% locked test. All snapshots of a launch land in the same split. Test launches are never written into a training dataset. Once enough of them have observed outcomes (at least 30 snapshots), they are frozen into a locked test set with a content hash. Benchmarks are scored only on launches that are active at the snapshot (at least 10 trades so far and some volume in the last 5 minutes — known at that moment). About three quarters of launches stop trading within minutes, and “a dead token stays dead” is too easy to count as skill, and every model version is scored on the same set. When a larger set is locked later, every earlier version is re-scored on it so the history stays comparable.

Scores are AUC for yes/no questions (0.5 = coin flip) and balanced accuracy for trajectories (≈33% = chance), with 90% bootstrap ranges computed over launches. There is no single “AI score”. A new version is deployed only if its average across the benchmarks it shares with the live version is at least as good; otherwise the old one stays live and the regression is recorded. Failed runs and worse versions are kept forever.

Datasets, benchmark sets, ledger entries, benchmark results and finished versions can't be edited or deleted; the database enforces it. Given the same dataset, feature version, configuration and seed, training produces a byte-identical model file.

How creator rewards fund compute

The TRAIN token earns creator rewards from trading on pump.fun. Rewards are recorded in the ledger only after their on-chain transaction is verified. They accumulate in the training treasury. When the treasury reaches the training target ($50.00 currently) and enough new outcomes have been observed, one training generation starts: compute is purchased, a new version is trained and benchmarked, and the result is published — better or worse.

Rewards don't trigger training trade by trade; they are batched into generations so each run is meaningful and each cost is visible.

Why everything is a probability

Memecoin launches are noisy. Two launches that look identical at 15 minutes regularly end in opposite places. The honest output is a base rate for setups like this one, with the size of the sample it came from. TRAIN reports those and labels them; it does not promise outcomes.

What is real right now

Treasury
None yet — the token hasn't launched; no rewards, no spending
What triggers training
New real data: enough newly observed launch outcomes since the last dataset
Compute
TRAIN's training server (CPU) — real training, nothing purchased yet
Explanation layer
Claude writes the wording; every number must come from TRAIN's model, or the answer is discarded
Market data
Real: GeckoTerminal (trades, prices) and pump.fun (creators, launches, curve status)
Model training & benchmarks
Real, on the real collected dataset

See the ledger, training runs and activity log.