Use this checklist after a simulated AI trading agent has enough paper evidence to review. The goal is to decide whether to keep testing, tighten one rule, reduce simulated risk, or retire the workflow, not to approve live execution.
Trading Boy does not execute live trades, hold funds, or provide financial advice. This checklist evaluates simulated paper-trading evidence only. It cannot prove live fills, future returns, emotional discipline, or suitability for real capital.
Evaluation checklist
Score each row before changing an agent prompt, persona, risk rule, or watchlist. A failed row should usually create another paper test, not a broad rewrite.
Gate
Question
Pass evidence
Stop condition
Version evidence
Can the sample be tied to one prompt, persona, rule, and market frame?
Every decision has a version label and review window.
The sample mixes unlabelled prompt or rule changes.
Rule fit
Did entries and skips follow the written setup and invalidation rules?
Journal notes cite the rule before the outcome is known.
The agent explains decisions after seeing the result.
Risk behavior
Did paper size, drawdown, and exposure stay inside the written limit?
Risk checks pass on winners and losers.
Paper gains depend on breaking the risk rule.
Skip discipline
Did the agent document when it did nothing?
Skips include the blocked condition and next review question.
Only entries are logged, so discipline is invisible.
Journal quality
Can a reviewer audit thesis, invalidation, result, mistake tag, and next action?
The journal contains structured fields for each decision.
The record contains confident summaries without evidence.
Sample quality
Is the sample large enough and complete enough to inspect?
Entries, exits, skips, misses, dates, and exclusions are visible.
Only selected winners or exciting trades are included.
Change decision
Is the next action narrow and testable?
The review chooses one change or more paper collection.
The review rewrites many variables from one small sample.
Scoring rubric for the review meeting
A checklist is only useful if the reviewer can turn it into a conservative next action. Use a simple score so the result does not become a vague debate about whether the agent "felt good" during the sample.
Score
Meaning
Allowed next action
Pass
The evidence is complete, the paper decision followed the written rule, and the reviewer can reproduce the reasoning from saved records.
Keep collecting paper evidence or compare the same version across a new review window.
Needs another sample
The agent behaved reasonably, but the sample is too small, one market regime dominates, or the journal lacks enough skipped-trade evidence.
Run the same version longer without changing the prompt, persona, watchlist, or paper risk settings.
Needs one fix
A repeated problem appears in the same checklist row, such as unclear invalidation, late exits, stale data, or oversized simulated size.
Change one field, write the version note, and restart the paper sample.
Stop the workflow
The sample mixes versions, ignores risk limits, invents missing fields, or cannot separate paper output from live-trading language.
Retire or rebuild the workflow before more paper records are collected.
Example completed checklist
Sample: A simulated AI paper-trading agent recorded 42 decisions across four weeks. The version label, watchlist, and output format stayed stable. The journal includes 18 entries, 16 skips, 6 exits, and 2 missed-trade notes.
Checklist result: Version evidence passes. Skip discipline passes. Journal quality mostly passes. Risk behavior fails because three paper entries exceeded the stated size cap after the agent described confidence as high.
Decision: The agent is not promoted, discarded, or moved toward live execution. The next paper test uses the same watchlist and setup rule, but the output format must include a hard size cap field before the agent can log an entry. The team then reviews the next sample with the same checklist.
Use with the review process
Use this checklist after the simulated AI trading agent review process. The process defines when to review; the checklist defines what evidence counts as clean enough to compare.
When a checklist produces a rule change, record it with AI trading agent prompt versioning. The next sample should make it clear which behavior changed and which variables stayed fixed.
Pair the final review with paper-trading limitations so the result stays framed as simulated evidence.
Checklist outputs
A completed checklist should produce one conservative output. Acceptable outputs include keep collecting paper evidence, tighten one output field, reduce simulated size, add one skip condition, split the sample by market regime, or retire the paper workflow. Avoid labels like approved, ready, or safe because paper evidence cannot prove those claims.
The strongest output is often boring: no change yet. If the agent followed rules, respected paper risk, and logged complete evidence, but the sample is still small, more observation may be the best next action. Paper trading becomes less useful when every short sample produces a prompt rewrite.
Reviewer notes to keep
Keep the completed checklist beside the sample, not separate from it. A useful review note names the version, date range, market regime, number of entries, number of skips, number of excluded records, and the exact checklist row that failed. That makes the next review easier because a second reviewer can see whether the issue was risk behavior, incomplete output, thin sample size, or unclear setup language.
Do not turn the checklist into a performance badge. A passing checklist means the paper evidence is cleaner, not that the agent is profitable, safe, or ready for live capital. The next action should still be a paper-mode action: collect another sample, tighten one field, or compare the same version across a different review window.
When the sample is not ready to score
Some paper samples should be excluded before the checklist is scored. Excluding a weak sample is better than pretending it supports a decision.
Mixed versions: the prompt, persona, risk settings, or watchlist changed during the review window without a version note.
No skipped trades: the journal records entries but not the moments when the agent correctly did nothing.
Missing risk context: paper size, invalidation, exposure, or drawdown state is absent from the records.
Permission drift: the agent output starts asking for live execution, custody, or account permissions instead of staying inside the paper-agent permission boundary.
What should an AI paper trading agent evaluation checklist include?
It should include rule fit, prompt version, journal quality, risk behavior, skipped trades, sample size, drawdown, market regime, and the next paper-test decision.
Can this checklist approve a live trading agent?
No. The checklist reviews simulated evidence only. Trading Boy does not execute live trades, hold funds, or provide financial advice.
When should an AI paper agent be changed?
Change one rule only when repeated paper evidence identifies a specific behavior problem, such as unclear invalidation, oversized simulated risk, late entries, or missing skip discipline.