The safe way for a beginner to start AI-assisted trading is an eight-week program on paper money: pick one market, write one rule the agent can follow — a rule that fits in a single sentence — and audit every decision the agent makes, every week. Go live only after passing the checklist at the end of this guide, and then with 5–10% of the capital you eventually intend to commit, because live fills, fees, and real-money emotion all behave differently from the simulation. Nothing in the program is novel; it is the routine experienced traders apply to any new strategy. It is worth writing down because most beginners skip it, and skipping it is the most reliable way to lose money in your first months.
| Weeks | Focus | The question it answers | You pass when |
|---|---|---|---|
| 1–2 | Translation check | Does the agent read the rule the way you meant it? | Every action and every skip matches your own reading of the rule |
| 3–4 | Variance tolerance | Can you sit through a normal losing streak? | You watched a drawdown without touching the kill switch |
| 5–6 | Decision-quality audit | Did every trade follow the rule as written? | Four yeses on every trade in the log |
| 7–8 | Stress test | Does the rule hold up around a scheduled volatility event? | Behavior in the event window matched the rule — including the skips |
Recurring economic data. The releases are scheduled months in advance, a published forecaster consensus exists before every one, and the same drill repeats every cycle — which is exactly what you want while you are still learning to read an agent's log. The comparison:
| Market type | Event timing | Consensus to compare against | Friction for a beginner |
|---|---|---|---|
| Recurring economic data (CPI, jobs report, FOMC) | Scheduled months in advance | Published forecaster consensus before every release | Lowest — the same drill repeats every cycle |
| Elections and politics | Resolution date known; the news flow is not | Polling averages — noisy and contested | Higher — long stretches with no resolution, plus headline whipsaw |
| Individual stocks | Earnings scheduled; other news arrives at random | Analyst estimates | Higher — company-specific surprises come without warning |
| Crypto | Trades 24/7, no schedule | No comparable consensus figure | Highest — there is no quiet period to learn in |
Pick a single rule you can write in one sentence:
"Enter Kalshi CPI contracts when the implied probability disagrees with the published forecaster consensus by more than 8 percentage points, and the spread is under 3 cents. Position size: 1% of paper capital. Max 4 positions open at once. Daily loss cap: 5% of account value."
Work the numbers once so every threshold means something. On a $1,000 paper account, 1% sizing is $10 per position — at 42¢ a contract, about 24 contracts. The daily loss cap is $50, and the 4-position limit caps concurrent exposure near $40. The trigger itself: if the forecaster consensus implies a 30% chance of the outcome and the contract asks 42¢ (implied 42%), the gap is 12 percentage points — past the 8-point threshold — and a book quoted 40¢ bid / 42¢ ask has a 2¢ spread, under the 3¢ limit. Both conditions hold, so the rule fires. If any of those numbers surprised you, that is the point of doing the arithmetic before the agent does.
The first two weeks are not a profit test. They are a translation test: does the agent interpret your rule the way you meant it? Watch every action. If the agent entered when you didn't think the conditions were met, the rule has an ambiguity you need to fix. If it skipped when you thought it should have entered, same — fix the rule. A defensible rule has explicit thresholds; "trade smart on CPI" is not a rule, because there is no way to check whether the agent followed it.
By week 3 the rule's translation should be clean. Now you're learning whether you can stomach the strategy.
Every rule that holds up over a year has weeks where it loses. Work the math once. Take a rule with a 55% hit rate — 55 of every 100 closed trades end positive. The chance that any given run of 5 trades all end negative is 0.455 ≈ 1.8%, which across 100 trades works out to roughly one 5-trade losing streak, and a better-than-even chance of seeing at least one. A streak like that is not evidence the rule broke; it is what a 55% hit rate looks like. If you would override the agent at trade 4 of that streak, you do not yet have the temperament for this strategy — and the override would convert a normal losing streak into a locked-in loss.
There is a reason the override urge is that strong. Tversky and Kahneman's cumulative prospect theory estimates (Journal of Risk and Uncertainty, 1992) put the weight of a loss at roughly 2.25 times an equivalent gain — so trade 4 of a losing streak feels far worse than trades 1–3 of the recovery will feel good. Knowing the number does not remove the feeling. It tells you the feeling is not information.
Watch yourself, not just the agent. If you find yourself reaching for the kill switch every drawdown, the strategy is wrong for you. Pick something with smaller swings.
Open the agent's trade log. For every trade — profitable or not — answer four questions:
Roll the answers into one number: rule-conformant actions divided by total actions. If the agent acted 40 times over the two weeks and 38 actions followed the rule as written, decision quality was 95% — and the two exceptions are the most valuable lines in the log, because they mark exactly where your intent and the agent's reading diverge.
If every trade gets four yeses, you have a clean strategy. The P&L over 6 weeks doesn't matter as much as the discipline of the execution. A strategy that loses money cleanly is fixable; a strategy that makes money sloppily will eventually catastrophically lose.
The audit is also what keeps you from adding your own impulse trades on top of the rule. The classic evidence on what discretionary overactivity costs retail accounts is Barber and Odean's study of 66,465 households with retail brokerage accounts, 1991–1996 (Journal of Finance, 2000): the most active fifth of households earned 11.4% a year net of costs while the market returned 17.9%, with trading costs — chiefly spreads and commissions — driving the gap. The audit habit channels the itch to act into reading logs instead of placing orders.
By now, run the agent through at least one scheduled high-volatility event — a CPI release, an FOMC meeting, an election milestone. Volatility is when bad strategies reveal themselves. The trades you want to inspect most carefully are the ones taken in the 30 minutes around the event.
Spreads are the thing to watch. A contract quoted 2¢ wide all week can be quoted 8¢ wide in the minutes around the release — which makes a round trip at the quotes cost 4 times what your rule assumed. A rule with a 3¢ spread cap should simply stop firing in that window. Confirm in the log that it did: the skips during the event are as much a pass condition as the entries before it.
Then the harder questions. Did the agent get blown out? Did it size up because "the setup was strong"? Did it ignore the spread widening? Each of these is a sign the rule needs tightening before any live deployment.
If, after 8 weeks of paper, all of these are true, you can consider going live with 5–10% of the capital you intend to commit:
Believing paper profits mean live profits. They don't necessarily. Live introduces slippage, latency, real-money emotion. The first 4 weeks live is a separate calibration period.
Overriding during a drawdown. A normal losing streak survived stays a normal losing streak. A normal losing streak interrupted becomes a locked-in loss, plus whatever the rule would have done next.
Running multiple strategies at once. Multiplies risk surface, makes attribution impossible, trains you to be inattentive. One strategy until you've mastered it — and one agent for at least the first six months, for the same reason: with two or more running, no outcome in the account traces cleanly back to a decision you can read.
Trusting an agent that can be talked past its risk cap. If the LLM ever agrees that "this trade is special, let's exceed the cap," your caps don't exist. Caps belong in code, not in prompts. Switch platforms.
Going live at full size. First live deployment should be 5–10% of intended size. Live diverges from paper; you want to measure how on small money.
Start with 5–10% of the capital you eventually intend to commit. If you plan to run the strategy with $2,000, the first live deployment is $100–$200. The small size is for measurement, not decoration: the first four live weeks are a calibration period in which you compare live fills against what the paper simulator assumed — same rule, same market, now with real spreads, real latency, and the occasional partial fill. Scale up only after you have measured that divergence. A strategy that has never traded a real dollar has not yet earned real size.
One test covers most of it: open the agent's log, pick any trade from the last six weeks at random, and explain why the agent took it — without looking at the P&L column. If you can do that for every trade you pick, you understand the strategy. If you need the P&L to tell you whether a trade was good, you are guessing, and live markets charge for guessing faster than paper does.
Two supporting checks. First, you actually completed the eight weeks — including at least one losing streak you did not enjoy and one scheduled volatility event the agent handled according to the rule. Second, your platform enforces its limits in code. Ask the vendor one direct question: can any conversation with the model raise a cap? Anything other than a flat no means the caps are suggestions.
TraderBear ships with paper money on by default. Pick one market, one rule, and audit every decision. Going live takes a deliberate, multi-step opt-in — not a slider.
Adopt a bear →