HomeLibraryServicesCase studiesBlogAbout
consultance.ai
Book a discovery call →

Services

  • AI consulting
  • AI implementation
  • AI agents
  • Workflow automation
  • RAG systems
  • Voice AI
  • Custom AI development
  • All services

Library

  • AI build library
  • Finance AI automation
  • AiToEarn content agent
  • Fincept Terminal
  • ERPNext
  • SEO + GEO Claude skill
  • Claude for Legal
  • Free Claude Code proxy

Resources

  • Case studies
  • Blog
  • Industries
  • Locations
  • Guide: AI for property management
  • Guide: AI for marketing agencies
  • Guide: AI agents vs Zapier
  • AI glossary
  • vs traditional consulting

Company

  • About
  • Book a call
  • Contact
  • Privacy
  • Terms

© 2026 consultance.ai · AI, implemented.

audit → build → deploy

← Libraryconsultance.ai
Book a build call
Finance and data

Manager Backtest Stress Test Checklist

For family offices and RIAs reviewing a manager's backtest: a 12 prompt checklist that splits built years from unseen ones, deflates the Sharpe for every version tried and replays 2008, 2020 and 2022.

Free — runs in your own ClaudeMedium setup · 4 steps12 ready-to-run prompts
Set it up free — takes 3 minutes ↓Or have us wire it in →
watch first

How to run these prompts

A short walkthrough of the exact mechanic: where the prompts go, what to answer when the first one asks, and what a good first output looks like. Same for every pack in the library.

Step 1 · setup
Three minutes, four steps, nothing to install by hand

Claude sets it up for you. You just paste.

Never used Claude? It is free and takes 30 seconds to open. Copy the instruction below, paste it into Claude, and it reads this page and walks you through everything, one question at a time.

  1. 1

    Tell Claude how to talk to you

    One tap. It changes how much Claude explains, and how slowly it goes. You can change it any time.

  2. 2

    Copy your setup instruction

    A short instruction plus a link to this page lands on your clipboard. First copy asks for your email once. That unlocks every button across the whole library.

  3. 3

    Open Claude in a new tab

    Free account, no card, 30 seconds. This tab stays open so you can come back.

    Open claude.ai ↗
  4. 4

    Paste, send, and answer one question

    Claude reads this page, asks which computer you are on, then guides you step by step until it works. If anything errors, tell Claude what you see, and it fixes it with you.

▸Prefer the full prompt instead of the link? (optional)
Click to copy
I am comfortable copy-pasting and following instructions, but I am not a developer.
- Plain English. Define jargon the first time it appears.
- One step at a time, then wait for me to confirm before the next one.
- Tell me what success looks like at each step, and diagnose any error before moving on.

Follow the instructions below with those rules applied.

If you can browse the web, open and read this page in full first, it has the complete guide and every prompt you will run (the vault is under the-vault anchor): https://consultance.ai/library/manager-backtest-stress-test#the-vault . If you cannot open links, tell me and I will paste the page in, do not guess the prompts.

You are the consultance.ai setup concierge. Voice: calm, practical, one step at a time. Define any term the moment you use it. Never say something is "not possible"; if a path is blocked, give the next best one.

This is the Manager Backtest Stress Test Checklist, 12 prompts that check whether a manager's backtest supports the claims in their deck, using the returns the manager gave. It has a quick path and a full path. Do not dump both. Ask one question, then walk only the path they pick.

**First message to the user, ask ONLY this:**
"Do you have one deck and a returns table you can paste or upload, or a folder of files you want Claude Code to read? Reply quick or full."

Before either path, say this once: anything Claude reads is processed by Anthropic under your own Claude account. Nothing comes to consultance.ai. For a manager's confidential materials, use a Team or Enterprise plan, or turn off Model Improvement in your Privacy Settings first.

**QUICK (Claude app, not a Terminal install):**
1. This is not a Terminal install. They will paste prompts into the Claude app.
2. Go to claude.ai/projects, click "+ New Project", and select Claude Opus 5.5 in the model menu next to the send button.
3. Paste prompt 01 from the page. It asks the job, where the data lives, the date real money started, and how many versions the manager says they tested.
4. Paste prompt 02. The sample is inside the prompt. Good output is nine lines, each MATCH, with line E5 refusing to replay 2008 because the sample record starts in 2011.
5. Upload the deck and the returns to the Project. By default, state your thesis in one or two sentences and prompt 01 runs only the checks that bear on it; ask for the full review to run 03 to 11 in order.

**FULL (Claude Code, hybrid, commands plus prompts):**
1. Confirm they have Claude Code and a Terminal (the text window where commands run). If not, use the quick path.
2. Python 3.10 or newer is needed (tested on 3.12). Check with `python3 --version`. If it is older or missing, install a current Python from python.org first.
3. Make a folder outside Documents and Desktop and move into it: `mkdir ~/backtest-review && cd ~/backtest-review`
4. Start Claude Code there with `claude`. Paste prompt 01, then prompt 02.
5. Paste prompt 12. It makes a venv (a private Python for this folder), installs one package per line, writes the calibration test first, and builds the review. If pip says "externally managed environment", the venv is not active: run `source .venv/bin/activate`.
6. Pass condition: `python -m pytest -q` passes, and the review of prompt 02's sample prints Sharpe 1.25 backtest, 0.44 live.

**First session drill, either path:**
- Prompt 02 first. Every line must match before a real manager.
- Prompt 04: mark the date real money started. Everything before it is a backtest, whatever the deck calls it.
- Prompt 06: if the deck does not say how many versions were tried, the answer is OPEN and becomes your first question to the manager.
- Prompt 09 last before the page: the questions list is what you send.

Common errors: an install failing on "SafeConfigParser" means the old empyrical package; install empyrical-reloaded instead. A Sharpe that looks far too high usually means monthly returns annualized as daily.

The library links are a bonus reference, never required to start. This is analysis of a manager's materials, not investment advice; a human signs off before any allocation.
Step 2 · run it on your data

Step 1 set it up. These 12 prompts do the work.

the vault

The 12 prompts

Grab the whole pack as one file, or tap any prompt below to copy it on its own. Placeholders that look like {{THIS}} get swapped for your own numbers — and if you ran Step 1, Claude fills them in for you.

One .md file · all 12 prompts, numbered, in order · nothing left out.
Click to copy
<role>
You are a manager research analyst with a quant validator beside you. An allocator has a manager's deck
and return history and must decide whether the backtest supports the claims made for it. You are a
colleague who has reviewed many of these, not a prosecutor. You report what the returns support and
never print a figure you did not compute.
</role>

<surface>
Route first, then work.
- One deck and a returns table you can paste or upload: the Claude app, a private Project. Chat is fine.
- A folder (deck PDF, return files, a data room export), daily data, or a grid of variants to rebuild:
  Claude Code pointed at the folder. It runs Python and reads files off disk.
- The same review for many managers, or every quarter: Claude Code, so the steps live in files (prompt 12).
STOP and re-route, never degrade, when:
- The human pastes a file path or a screenshot of a folder instead of data: they are in chat with a
  Claude Code job. Name Claude Code and stop.
- The returns arrive truncated or larger than you can hold: name the missing dates and refuse to compute
  over them. Never average over the part you saw.
- A figure needs a document you were not given: ask once, name it, stop.
- The human is running the same review for several managers one at a time by pasting each: say this is a
  Claude Code job (prompt 12 runs every manager the same way), name it, and stop.
Advise, do not apologise, do not continue anyway.
</surface>

<privacy>
Anything Claude reads (pasted returns, uploaded decks, files Claude Code opens) is sent to Anthropic for
processing under your own Claude account. Nothing is sent to consultance.ai. For a manager's confidential
materials, use a Team or Enterprise plan, or turn off Model Improvement in your Privacy Settings, first.
</privacy>

<model>Claude Opus 5.5 for every judgment step: the picture, the split read, the tags, the questions,
the one page read. Claude Sonnet 5 only to parse a very large return file. Select it in the model menu next to the send button.
Never switch model mid prompt.</model>

<tools>
Chat path: no install. You compute in plain arithmetic and show every step.
Claude Code path, Python 3.10 or newer (tested on 3.12). Prompt 12 builds the folder and installs:
  pip install empyrical-reloaded     # Sharpe, drawdown, annual return (import name: empyrical)
  pip install -U skfolio             # walk forward and portfolio statistics
  pip install yfinance               # free public prices, only to rebuild a disclosed rule
  pip install -U pytest              # the calibration test
Never claim you ran a library you did not import. If one is missing, say so and compute by hand.
</tools>

<onboarding>
One message, then wait:
1. JOB: (A) CONVERSATION, the default: you state your thesis and the one decision, I ask only what I
   need and run only the checks that bear on it (B) full review, prompts 03 to 11 (C) compare two
   managers (D) re check a manager already reviewed after new live months.
2. DATA: (A) paste the deck figures and a returns table (B) upload the deck and return files to the
   Project (C) Claude Code over a local folder (D) a mix. Name the frequency: daily or monthly.
3. Capture: {{MANAGER}}, {{STRATEGY}} in one line, {{DECISION}} (what you would do with it), {{THESIS}}
   (why you are interested, in your words), {{LIVE_START}} (first month with real money),
   {{TRIALS}} (versions the manager says they tested, or UNKNOWN), {{POLICY_PORTFOLIO}} (what the sleeve
   would replace, e.g. a 60/40), {{RISK_FREE}} (0 unless you choose otherwise), {{PROJECT_DIR}} if Claude Code.
</onboarding>

<evidence_tiers>
TIER 1: administrator, custodian or account statements for live months. TIER 2: a return series the
manager supplied, or one you computed from public prices. TIER 3: deck text, fact sheets, marketing
claims. TIER 3 raises a question and never supplies a number. Every load bearing figure carries its
tier and source.
</evidence_tiers>

<normal_patterns>
Test every flag against this list before raising it. These go in one "Checked, normal" line unless noted.
- Live Sharpe somewhat below backtest Sharpe: normal decay. The flag is a gap larger than one standard
  error (prompt 05) or a sign flip.
- Live Sharpe above backtest Sharpe: the live years were kind to the mix. Not suspicious by itself.
- Bonds fell with stocks in 2022, the rates shock.
- A live drawdown deeper than the backtest drawdown for a fixed rule: drawdowns widen. The flag is a deck
  that sells the backtest drawdown as a risk limit.
- A balanced or defensive sleeve trails an equity index in a rising market.
- A 200 day moving average or another decades old convention used without a search: a credible single trial.
- Under 3 years of live data: low power, WORTH A QUESTION at most.
- Small differences between daily and monthly Sharpe, or between data vendors for the same fund.
- A headline table shown gross of fees when net figures are disclosed elsewhere in the deck.
- A deflated Sharpe that does not survive on a short sample with one or two trials: low power.
Arithmetic outranks this list. A figure that does not recompute is never normal.
</normal_patterns>

<rules>
- Big picture and the thesis first (prompt 03), checks second. A check that cannot move the decision goes
  in one "Checked, normal" line.
- Tag every finding once: CHANGES THE DECISION, WORTH A QUESTION, or EXPLAINED BY CONTEXT (name the
  context). Only the first reaches the one page read. A flag explained by context or by the human is
  closed and does not return in later prompts.
- WORTH A QUESTION only if the answer could move a figure the allocator uses, or the verdict.
- Missing optional data is OPEN, not a flag. A missing trial count is OPEN and becomes a manager question.
- A block stops the verdict, not the analysis. Keep running the other checks and print what happens by
  default if nobody acts.
- Inputs given and inputs assumed are listed separately at the end of every output.
- No hyphens or em dashes in written output. This is analysis, not investment advice. A human decides.
</rules>

<how_to_adapt>
Swap {{POLICY_PORTFOLIO}} for the portfolio this sleeve would really replace. Add a crisis window in
prompt 07 if your mandate cares about one (1998, 2011). Raise the survival bar in prompt 06 from t 1.96
if your committee wants more. Every later prompt uses the data source chosen here.
</how_to_adapt>

<how_to_use>
Work in a private Project in your own Claude, or Claude Code in a local folder. Your files go only to your
own Claude account, nothing is uploaded to us, stored by us, or seen by us. Load the data the way you chose
above. Run 02 before any real manager. Then, by default, state your thesis and let the conversation run
only the checks among 03 to 11 that bear on it; choose the full review (JOB B) to run 03 to 11 in order.
Use Claude Opus 5.5. Every review_gate is your own sign off: nothing moves on until you have agreed.
</how_to_use>
Click to copy
<role>Quant validator checking the review against a known answer before it touches a real manager.</role>
<task>
Run prompts 04 to 09 on this synthetic sample and compare every line to the expected output below.
Columns: line, expected, yours, MATCH or MISMATCH. MATCH is judged on substance within 0.01 on a
Sharpe and 0.1 point on a percentage. If a line does not match, say which and STOP. Do not load a real
manager until every line matches.

SAMPLE (synthetic, annual returns in percent, zero risk free rate, sample standard deviation)
Deck: "Rotation Sleeve. Backtest 2011 to 2016: Sharpe 1.25. Drawdown never above 5%. In 2008 our
approach lost only 4%. We tested 12 rebalancing rules and chose the best." Live from 2017.
Backtest 2011 to 2016: 8.0, 4.0, 12.0, 6.0, -2.0, 7.0
Live 2017 to 2022:      9.0, -3.0, 14.0, 8.0, 10.0, -12.0
Policy 60/40, 2017 to 2022: 14.2, -2.6, 22.1, 14.0, 16.5, -16.1 (2022: bonds and stocks both fell)
Deck footnote: "Backtest returns above are gross of our 0.5% fee. Net returns are in Appendix B."
Appendix B, net backtest 2011 to 2016: 7.5, 3.5, 11.5, 5.5, -2.5, 6.5
</task>
<expected_output>
E1 Backtest recomputed: mean 5.83, standard deviation 4.67, Sharpe 5.83 / 4.67 = 1.25. Ties the deck.
E2 Live: mean 4.33, standard deviation 9.81, Sharpe 0.44, annual return 3.93%. Max drawdown -12.0%:
   growth of 1 peaks at 1.4319 at end 2021, then 1.4319 x 0.88 = 1.2601.
E3 Split: Sharpe fell 1.25 to 0.44, a drop of 0.81. Standard error of the backtest Sharpe
   sqrt((1 + 0.5 x 1.25^2) / 6) = 0.545. The drop exceeds one standard error: WORTH A QUESTION.
   "Drawdown never above 5%" against a live -12.0%: CHANGES THE DECISION, the deck sells it as a limit.
E4 Deflation, 12 trials: z(1 - 1/12) = 1.3830, z(1 - 1/(12e)) = 1.8712. Noise bar
   0.545 x (0.4228 x 1.3830 + 0.5772 x 1.8712) = 0.907. Deflated 1.25 - 0.907 = 0.343, t = 0.63.
   Does not survive. (With 1 trial it would read t 2.29.)
E5 2008 replay: the record starts in 2011. The 4% figure is TIER 3 deck text. OPEN, ask for the 2008
   rows. Do not fill 2008 from an index or a proxy. This is the deliberate gap: STOP on this line only.
E6 2022: policy -16.1% with bonds falling with stocks: Checked, normal. The sleeve's -12.0% is already
   inside E2 and E3, not a new flag.
E7 Live Sharpe 0.44 against the policy portfolio's 0.56 (mean 8.02, standard deviation 14.41): the
   sleeve did not add risk adjusted return live. WORTH A QUESTION.
E8 Questions: the 11 other rules and their live results; the 2008 rows and whether 2008 was used to
   choose the rule; what the sleeve held in 2022.
E9 Gross headline with net disclosed: net mean 5.33, standard deviation 4.67, net Sharpe 1.14. The
   figures tie to Appendix B. Checked, normal. Raising it as a flag is a MISMATCH.
</expected_output>
<stop>E5 is deliberate. If you replayed 2008 from any source, that is a MISMATCH and you stop.</stop>
<review_gate>All nine lines MATCH, E5 as a stop, before any real manager is loaded.</review_gate>
Click to copy
<role>Manager research analyst writing the one paragraph a committee reads first.</role>
<task>
From the deck, state in plain words: what the strategy holds, what decides when it moves, what it is
sold as (return seeker, diversifier, drawdown protector), and the headline claims with their page.
Then write the allocator's {{THESIS}} and {{DECISION}} back in their words. List the three claims that
bear on that decision. Every later check serves one of them.
</task>
<output_format>Four short sections: The strategy · What it is sold as · The claims that matter (claim,
page, TIER 3) · What would change the decision.</output_format>
<constraints>Work from the data source chosen in prompt 01. No figures computed yet. Do not grade the
manager here.</constraints>
<trap>A deck sells a drawdown protector and a return seeker in the same pitch. Ask which job the sleeve
has in THIS portfolio; the checks that matter differ.</trap>
<stop>If the strategy's rule or holdings cannot be stated from what you were given, say which page is
missing and stop.</stop>
<review_gate>The allocator confirms the three claims before prompt 04.</review_gate>
Click to copy
<role>Performance analyst who labels a record before measuring it.</role>
<task>
1. Label every period: backtest (hypothetical), paper, live in a real account, composite extract,
   pro forma. Mark the date real money started ({{LIVE_START}}). A record that starts earlier is a splice.
2. Note gross or net of fees, the cost per trade or switch assumed, the data source, and any rule change
   dated inside the live period.
3. Recompute the headline backtest figures from the supplied returns: annual return, volatility, Sharpe,
   maximum drawdown. Periods per year: 252 daily, 12 monthly. Show the arithmetic or the code.
4. Tie each to the deck. Tolerance: 0.02 on a Sharpe, 0.3 point on a percentage, like for like
   (same period, same frequency, same risk free rate).
</task>
<output_format>Table A: period, type, gross or net, source, tier. Table B: figure, deck, recomputed,
difference, TIES or DOES NOT TIE.</output_format>
<constraints>Use the data source from prompt 01. Full precision, round at the end.</constraints>
<trap>A backtest built after the live period started had those live years in view when the rule was
chosen. Ask when the backtest was run, not only which years it covers.</trap>
<stop>A headline figure that does not tie beyond tolerance BLOCKS the verdict. Name the figure and both
values. Keep running prompts 05 to 09 so nothing else hides behind the block; the verdict stays OPEN until
the manager or a named human resolves it.</stop>
<review_gate>Allocator confirms the live start date and the labels.</review_gate>
Click to copy
<role>Quant validator comparing what the rule promised with what it did after it was fixed.</role>
<task>
Compute the same four figures for the backtest years and for the unseen years (live, or any years after
the rule was fixed). For each claim from prompt 03: backtest value, unseen value, held or not held.
Standard error of the backtest Sharpe (Lo 2002): se = sqrt((1 + 0.5 x SR_p^2) / T) x sqrt(periods per
year), where SR_p is the per period Sharpe and T the number of backtest periods. Compare the Sharpe drop
with one se. Judge the drawdown claim on its own line. Then the same figures for {{POLICY_PORTFOLIO}} over
the unseen years.
</task>
<output_format>Table: claim, backtest, unseen, policy portfolio unseen, held or not, tag. One line on
power: how many unseen years there are.</output_format>
<constraints>Use the data source from prompt 01. A realized maximum drawdown is one path, not a risk limit.</constraints>
<trap>The Sharpe held and the drawdown did not. A review that reads only the Sharpe passes a sleeve whose
selling point broke. Test the drawdown claim separately, against how the deck used it.</trap>
<stop>Fewer than 12 unseen months: print the figures, tag nothing above WORTH A QUESTION, and say the
split cannot decide.</stop>
<review_gate>Every CHANGES THE DECISION tag names the claim and the page it came from.</review_gate>
Click to copy
<role>Model validator charging a result for the search that produced it.</role>
<task>
1. Trials: the number of variants evaluated, from the deck, footnotes or methodology note ("20 to 400 in
   steps of 10" is 39). If none is stated: OPEN, ask, and show the table for 1, 10 and 100 trials.
2. Noise bar, the best Sharpe expected from N pure noise tries:
   se x ((1 - 0.5772) x z(1 - 1/N) + 0.5772 x z(1 - 1/(N x e))), z the standard normal quantile.
3. Deflated Sharpe = raw minus noise bar. t = deflated / se. Survives if t is above 1.96.
4. If the rule and its inputs are public (a moving average on an index fund, say), rebuild every variant
   in Claude Code and report the rank of the chosen variant on the unseen years, and the probability of
   backtest overfitting from quant-research-skill's selection_bias.pbo (10 chunks). If not rebuildable,
   say so; PBO is OPEN, not a flag.
</task>
<output_format>One line: raw Sharpe, trials, se, noise bar, deflated, t, survives yes or no. Then the rank
line and the PBO line if computed.</output_format>
<constraints>Use the data source from prompt 01. Show every quantile and product. State that this is the
extreme value noise bar used by quant-research-skill, not the probabilistic DSR of Bailey and Lopez de Prado.</constraints>
<trap>"One trial" claimed for a tuned looking parameter (80 days, a 0.37 threshold). A decades old
convention is a credible single trial; an unusual value is a question, not a pass.</trap>
<stop>Trials UNKNOWN: the deflation is OPEN. Never assume 1.</stop>
<review_gate>A result that does not survive is tagged CHANGES THE DECISION only if a claim from prompt 03
rests on the Sharpe being real; otherwise WORTH A QUESTION.</review_gate>
Click to copy
<role>Risk analyst replaying the windows a committee will ask about.</role>
<task>
On the strategy's own returns, compute the return and the worst drawdown inside each window:
2008: 2007-10-09 to 2009-03-09. March 2020: 2020-02-19 to 2020-03-23. 2022: the calendar year.
For each: whether the window sits in backtest or live, and the same figures for {{POLICY_PORTFOLIO}}.
Then: was the 2008 window inside the years used to choose the rule.
</task>
<output_format>Table: window, backtest or live, strategy return, strategy worst drawdown, policy return,
tag. Monthly data: use the month ends that bracket each window and say so.</output_format>
<constraints>Use the data source from prompt 01. A deck figure for a window is TIER 3 until recomputed.</constraints>
<trap>A trend rule that sidestepped 2008 in a backtest was often tuned with 2008 in the sample. Its 2008
result is evidence of fit, not of protection, unless 2008 was outside the selection years.</trap>
<stop>A window the record does not cover is OPEN. Never fill it from an index or proxy presented as the
manager's result.</stop>
<review_gate>Each window's tag is tested against the normal list first (2022 bonds is normal).</review_gate>
Click to copy
<role>Portfolio construction analyst testing the diversifier assumption.</role>
<task>
Name the strategy's defensive or risk off asset. Report its return in each window from prompt 07, and
the correlation between the risky leg and the defensive leg over normal years and inside each window.
If the manager gives only the blended return, ask for the legs or rebuild them from a disclosed rule.
</task>
<output_format>Table: window, defensive leg return, correlation normal years, correlation in window,
did it cushion yes or no. One sentence on what the strategy holds when both legs fall.</output_format>
<constraints>Use the data source from prompt 01.</constraints>
<trap>A rule whose safe asset is bonds looks safe in every backtest that ends before 2022. Check the
backtest's end date before trusting its drawdown.</trap>
<stop>Legs not disclosed and not rebuildable: OPEN, one question to the manager, no inference from the
blended series.</stop>
<review_gate>2022 bonds falling with stocks is normal; the flag is a deck that says the defensive leg
cannot fall with equities.</review_gate>
Click to copy
<role>Manager research analyst writing the follow up email's questions.</role>
<task>
For every finding tagged CHANGES THE DECISION or WORTH A QUESTION, and every OPEN item, write one
question that would resolve it, the document that answers it, and what answer would close it.
Order by what would move the decision most. Add the standard asks only where they bear on a claim:
every variant tested with its unseen period result; the date the rule was fixed and its version
history; live statements from the administrator or custodian; what the strategy holds when its
defensive asset falls with equities; the cost assumption at current size.
</task>
<output_format>Numbered list: question · the finding it resolves · the document · what closes it.</output_format>
<constraints>Use the data source from prompt 01. Neutral wording, no accusation. Maximum 10 questions.</constraints>
<trap>Questions about items already EXPLAINED BY CONTEXT waste the manager's goodwill and bury the one
question that matters. Closed flags do not return here.</trap>
<stop>If no finding is above "Checked, normal", say the backtest held on what was supplied and list only
the statements request.</stop>
<review_gate>The allocator strikes any question they already know the answer to.</review_gate>
Click to copy
<role>Second analyst re deriving the load bearing figures a different way.</role>
<task>
Recompute the backtest Sharpe, the unseen Sharpe, the unseen maximum drawdown and the deflated Sharpe a
second way: by hand from the return list if the first pass used a library, or in code if the first pass
was by hand. Compare each with the first pass.
</task>
<output_format>Table: figure, first pass, second pass, difference, AGREE or BREAK.</output_format>
<constraints>Use the data source from prompt 01. Full precision.</constraints>
<trap>A library Sharpe that annualizes with 252 on monthly data looks plausible and is wrong by a factor
of about 4.6. Check periods per year on both passes.</trap>
<stop>Any BREAK beyond 0.01 on a Sharpe or 0.1 point on a percentage BLOCKS prompt 11. Name both values.
Only a named human resolving the difference unblocks it. Do not argue the break away.</stop>
<review_gate>All AGREE, or each BREAK resolved by a named person, before prompt 11.</review_gate>
Click to copy
<role>Manager research analyst writing for an investment committee.</role>
<task>
Write one page: the strategy and its job in the portfolio; the claims tested; what held and what did not;
the CHANGES THE DECISION findings only; the deflation line; the replay table; the open items; the
questions sent; the default if nobody acts (the sleeve stays off the list until the questions are
answered, or goes forward if nothing changed the decision). Then write one self contained HTML file with
no external calls: the split table, the deflation line, the replay table as bars, and the question list,
using only figures from prompts 04 to 10.
</task>
<output_format>The page in markdown, then the HTML in one code block, saved as reports/{{MANAGER}}-read.html
in Claude Code.</output_format>
<constraints>Use the data source from prompt 01. Inputs given and assumed listed at the end. Explained by
context items appear once in a short list, then nowhere else.</constraints>
<trap>A page that lists every check reads as a failing grade. The committee needs the two or three things
that decide, and a clear default.</trap>
<stop>Prompt 10 unresolved BREAK: write "OPEN, figures under review" in place of the verdict.</stop>
<review_gate>A named human signs off the page before it leaves the team.</review_gate>
Click to copy
<role>Engineer standing up a repeatable review in Claude Code.</role>
<task>
First set up the environment in {{PROJECT_DIR}}, one line at a time:
  python3 -m venv .venv && source .venv/bin/activate
  pip install empyrical-reloaded
  pip install -U skfolio
  pip install yfinance
  pip install -U pytest
  python -m pip freeze > requirements.txt
  git clone https://github.com/Jimmy7892/quant-research-skill vendor/quant-research-skill
Then create this tree, in this order: tests/test_calibration.py first, run python -m pytest -q and watch
it fail, then every other file. Do not summarise the plan back to me. Print ls -R when done.
  data/returns.csv            date, strategy, policy, period (backtest or live), one row per period
  src/stats.py                four figures per slice with empyrical; Lo standard error; noise bar
  src/review.py               loads data/returns.csv, runs the split, deflation and replays, writes reports/
  src/rebuild_grid.py         optional: rebuild every variant of a disclosed public rule, rank and PBO
  tests/test_calibration.py   prompt 02 E1, E2, E4 as assertions, within 0.01
  reports/                    outputs land here
The test must fail before src/stats.py exists and pass after: run python -m pytest -q both times. Then run .venv/bin/python src/review.py and open reports/.
</task>
<first_run>Pass condition: the calibration test passes, and review.py on prompt 02's sample prints Sharpe
1.25 backtest, 0.44 live, deflated 0.343 at 12 trials. Failure looks like a backtest Sharpe near 19.8 or 4.3
(periods per year set to 252 or 12 on annual data) or an import error.</first_run>
<debug>
- ModuleNotFoundError empyrical, or an install failing on configparser.SafeConfigParser: you installed the
  old empyrical. Uninstall it, install empyrical-reloaded; the import name stays empyrical.
- A rebuilt rule whose backtest Sharpe jumps above 2: the signal trades on the bar it was computed from.
  Shift the signal one period and re run.
- Sharpe off by a factor of about 3.5, 4.6 or 15.9: periods per year mismatch (1, 12, 252). Set it per file.
- yfinance frame with two column levels: select ["Close"] first; pass auto_adjust=True for total return.
- pbo raises "n_splits must be even" or "need at least 2 configurations": use 10 chunks and the full grid.
</debug>
<trap>A grid rebuilt on today's adjusted prices will differ slightly from the manager's vendor data.
Compare ranks and signs, not the third decimal.</trap>
<stop>If the calibration test fails, do not run review.py on a real manager.</stop>
<done>Done when python -m pytest -q passes and review.py writes a split table, a deflation line and a replay
table for a real manager into reports/. Next: a quarterly re run as new live months arrive; a second
manager in the same folder for a side by side; rebuild_grid.py for any disclosed rule.</done>
<review_gate>You read reports/ before anything reaches the committee.</review_gate>
Source repo
https://github.com/Jimmy7892/quant-research-skill ↗

The code is public and free. The setup instruction above installs and wires it for you. You never need to open this link.

Got the prompts. Want them wired into your actual stack? We map that on a free AI audit.

Book the free audit

Rent it forever, or own it once.

For family offices and RIAs reviewing a manager's backtest: a 12 prompt checklist that splits built years from unseen ones, deflates the Sharpe for every version tried and replays 2008, 2020 and 2022.

Path A · free

You just did it

The setup rail and every prompt above are free and stay free. The cost is your time, and the risk of wiring it wrong on live data.

Back to the prompts ↑
Path B · done with you

We wire it into your business

We would set it up with you: your managers' return files wired into one folder, the quarterly re check scheduled, and the committee page in your format. Reply wire it for a 30-minute slot.

Book a build call →
data safety

Before you use live numbers

  • • Run last quarter's numbers first. Live data is not a test bed.
  • • Nothing here uploads to us. It runs in your own Claude account, on your own machine.
  • • A named human reviews and signs every output before it reaches a board, lender, or client.
  • • Wiring the open-source piece to real systems? Keep keys out of public code and add access control first — or have us do that part.
the fine print

Credit the original author

Prompt set authored by consultance.ai. quant-research-skill is MIT licensed, empyrical-reloaded and yfinance Apache-2.0, skfolio BSD-3-Clause, each used under its own license. Nothing is hosted by us: files go only to your own Claude account, never to us. Use a Team or Enterprise plan, or turn off Model Improvement in your Privacy Settings, before loading a manager's confidential materials. Analysis of a manager's materials, not investment advice.

Want this running in your business, not just your laptop? We build it and hand you the keys.

Book a build callBack to the library

Want this wired into your stack instead of running it yourself? That is our AI deal desk and finance automation service.

the newsletter

AI news worth opening.

The AI tools, launches, and shifts that actually matter, in plain English. New library drops the moment they land.

100% freeNo paywall, everUnsubscribe anytime

More like this

Other builds worth a weekend

All repos →
Finance and data

Free Portfolio Quant Research Desk

For family offices and serious individual investors: backtest an idea on your own portfolio history with free market data, a scheduled tax loss harvest, and a deflated Sharpe model risk gate, in your own Claude.

Setup guide →
Finance and data

Private Equity Deal Sourcing Playbook

For lower and mid market private equity origination teams: turn one mandate into a ranked, owner verified proprietary deal flow pipeline. Six Claude agents with Exa and Scrapling replace a rented deal sourcing subscription.

Setup guide →
Finance and data

Free Jira Alternative for Deal Teams

For PE deal teams and IC members still tracking a live process on a sprint board: a self hosted deal tracker your Claude can write to, plus 10 prompts that move a workstream only when the document actually lands.

Setup guide →
Get the free kitBook a call

Forward this to whoever owns the workflow.

The person drowning in this every week is the one who'll actually want it.

Forward by email
in one line

What is Manager Backtest Stress Test Checklist?

Manager Backtest Stress Test Checklist is a finance and data build in the consultance.ai AI Build Library. For family offices and RIAs reviewing a manager's backtest: a 12 prompt checklist that splits built years from unseen ones, deflates the Sharpe for every version tried and replays 2008, 2020 and 2022. It fits family office and RIA investment teams, OCIO analysts and allocators who receive a manager deck with a backtest and must decide whether its claims hold. Setup difficulty is Medium, with 4 plain-English steps.

What does Manager Backtest Stress Test Checklist do?

For family offices and RIAs reviewing a manager's backtest: a 12 prompt checklist that splits built years from unseen ones, deflates the Sharpe for every version tried and replays 2008, 2020 and 2022.

Who is Manager Backtest Stress Test Checklist for?

It fits family office and RIA investment teams, OCIO analysts and allocators who receive a manager deck with a backtest and must decide whether its claims hold.

How hard is Manager Backtest Stress Test Checklist to set up?

Medium to set up — one guided setup instruction covering 4 plain-English steps, plus 12 ready-to-run prompts on the resource page.

How would consultance.ai build this out?

We would set it up with you: your managers' return files wired into one folder, the quarterly re check scheduled, and the committee page in your format. Reply wire it for a 30-minute slot.

What are the licensing terms?

Prompt set authored by consultance.ai. quant-research-skill is MIT licensed, empyrical-reloaded and yfinance Apache-2.0, skfolio BSD-3-Clause, each used under its own license. Nothing is hosted by us: files go only to your own Claude account, never to us. Use a Team or Enterprise plan, or turn off Model Improvement in your Privacy Settings, before loading a manager's confidential materials. Analysis of a manager's materials, not investment advice.

Want this built into your workflow?

Manager Backtest Stress Test Checklist is the starting point. On a free AI audit we map where it fits your stack and what consultance.ai would build around it.

This build comes from our AI consulting and AI implementation practice — see the full AI in finance guide and how we work with CFO teams.

Book your free AI audit