Part XV · Chapter 79

The Backtesting Mindset

First published 25 Aug 2026 · Last verified 29 Aug 2026

Keshav Raj Pyakurel spent thirty-one years reading numbers for a living. Not writing them, not selling them — reading them, the way a mechanic reads an engine noise before it becomes a breakdown. Most of that career was spent inside Nepal Rastra Bank's supervision wing, going through the balance sheets of commercial banks and development banks line by line, flagging the ones whose capital cushions were thinner than their public filings suggested. He was the kind of officer who kept his own private ledger of every bank he supervised, cross-checked against the official file, because he did not fully trust any single source — including, by his own admission, himself. He retired in the middle of 2021, at almost exactly the moment NEPSE was setting the all-time high it would not revisit for years. For about four months he did what half of Kathmandu was doing that year — he bought whatever was moving, felt briefly brilliant, and then watched a fair portion of his retirement gains evaporate through the back half of 2021 and into 2022. It embarrassed him enough that he stopped trading almost entirely for a year and started reading instead.

By the time he found the Canon Score — the seven-dimension, hundred-point scoring framework this book built across Part XIII, covering things like earnings quality, capital adequacy, promoter behaviour, and liquidity — he had already rebuilt his old habit of keeping a private ledger, except now the ledger tracked every stock he scored and every decision he made using that score. A year of this felt reassuring. It also felt, to a man trained to distrust reassurance, insufficient. "The score has made sense to me for twelve months," he told a former colleague over tea in Battisputali. "That is not evidence. That is a good feeling with a spreadsheet attached." He wanted to know something harder: if he had been using this score for the last seven or eight years, not just the last one, would it actually have kept him out of trouble and put him into the winners — or would it have quietly failed in ways a single good year could never reveal? That question is what backtesting is for, and it is what this chapter, and the whole of Part XV, is built around.

Lesson 79.1 — What Backtesting Actually Means

Backtesting is a simple idea wearing an intimidating name. It means taking a rule — any rule, whether it is the full Canon Score, a single entry signal like "buy when a bank trades below book value with a clean NPL trend," or an exit rule like "sell if a stock breaches its 90-day low on above-average volume" — and running that rule mechanically against data from the past to see what would have happened if you had actually followed it, trade by trade, at the time.

That last phrase, at the time, is the entire discipline. It is the difference between backtesting and the two things retail investors usually do instead and mistake for backtesting.

The first substitute is "the strategy feels right." Keshav had this in spades. The Canon Score felt right to him because every time he scored a stock highly, he could construct a story afterward for why it made sense — strong promoter holding, low debt, decent dividend history. The trouble is that a plausible story after the fact is not a test. A story can be built around almost any outcome, good or bad; that is what stories are for. A backtest, by contrast, commits to the rule before it sees the ending and then simply records what happened. It removes your permission to explain away the losses.

The second substitute is "it worked for my friend." This is social proof standing in for evidence. Keshav's nephew had made good money in Api Power and a hydropower IPO the year before and was certain that buying any hydropower counter within a month of monsoon onset was a reliable strategy. One profitable friend is a sample size of one, observed with hindsight, filtered through memory that tends to keep the wins and quietly discard the losses. It is not data; it is an anecdote wearing data's clothes.

KEY CONCEPT Backtesting means applying a precisely defined rule to historical data, trade by trade, using only the information that would genuinely have been available on each decision date — then honestly recording every result, not just the flattering ones.

What made Keshav a natural at this, once he understood what was being asked of him, was that his old NRB job had trained him to do something structurally identical: he used to take a bank's lending rule (say, a loan-to-value ceiling) and check it against years of actual loan files to see whether the rule had been followed and whether following it had actually protected the bank. Backtesting a trading or scoring rule is the same exercise, just aimed at a stock's price and volume history instead of a bank's loan book. The rule is the hypothesis. The historical data is the evidence. The backtest is the honest cross-examination of the two.

He picked a first, deliberately small test to build the habit: the exit rule from Part XII — sell a position if daily traded volume falls below a set multiple of the position size for more than five consecutive sessions (a liquidity-based exit meant to get you out before you become the seller nobody wants). He pulled five years of daily volume data for a dozen mid-cap manufacturing and hotel counters he already followed, and for each one asked a simple question on each trading day: given only what was known up to and including that day, would this rule have told me to sell? Not what he remembered feeling on that day. Not what the stock did afterward. Only the rule, applied mechanically, to the data as it stood at that moment.

Lesson 79.2 — The Three Biases That Quietly Ruin Amateur Backtests

Keshav's first attempt at a fuller backtest — running the Canon Score against eight years of data across roughly forty NEPSE counters — produced a hit rate so flattering he did not believe it. He was right not to. He had, without realising it, walked into the three failure modes that ruin almost every amateur backtest, in NEPSE or anywhere else.

The first is survivorship bias. Keshav's stock list was built from today's NEPSE company list — the ~260 or so companies currently trading. But NEPSE eight years ago had a meaningfully different roster: dozens of development banks and finance companies that existed then have since disappeared, mostly through forced mergers driven by Nepal Rastra Bank's 2015 directive raising commercial banks' minimum paid-up capital fourfold, to Rs 8 billion, with smaller BFIs facing their own consolidation pressure. Keshav actually remembered this directive intimately — he had helped enforce it from inside NRB. Companies that failed to raise capital, or whose books turned out weaker than advertised, were absorbed into larger institutions and vanished from the ticker. If your backtest only includes companies that are still independently listed today, you have quietly deleted every failure from the sample. The past looks safer than it was, because the stocks that did not survive are not there to drag the average down.

WARNING A backtest built only from companies currently listed on NEPSE silently excludes every company that failed, merged, or was delisted along the way — which flatters every "buy and hold" style rule tested against it.

The second is look-ahead bias — accidentally feeding the rule information it could not have had at decision time. Keshav's Canon Score uses full-year audited earnings per share as one of its inputs. When he first ran his backtest, he pulled EPS from annual reports and applied that full-year audited figure to score a stock as of, say, mid-Poush (roughly December–January) of that same fiscal year. But audited annual results for a NEPSE company are typically published months after the fiscal year closes — often not until Ashoj or Kartik of the following year. In mid-Poush, in real time, an investor would have had only the unaudited quarterly figures released so far, which can differ meaningfully from the final audited number once provisioning, write-offs, or restatements come through. By plugging in the audited figure too early, Keshav's backtest was quietly cheating — giving his rule knowledge from the future.

The third, and the one that took him longest to see in his own work, is overfitting. After noticing the audited-EPS problem, Keshav started adjusting the Canon Score's internal weights — nudging the liquidity dimension up a little, the promoter-holding dimension down a little — until the backtest's returns looked even better. Each adjustment was justified by some specific stock in his sample. After two weekends of this he had a version of the score that explained his forty-stock, eight-year history almost perfectly. It also, he began to suspect, explained nothing at all. He had not found a better rule; he had sculpted a rule around the exact bumps and dips of one particular, noisy, forty-stock sample. Give that same tuned rule a different set of stocks, or the same stocks over a different stretch of years, and there is no reason to expect it to perform anywhere near as well — because it was never describing a real, repeatable relationship. It was describing coincidence, dressed up to look like insight.

CASE IN POINT Keshav tuned the Canon Score's weights until it "explained" eight years of data on forty stocks almost perfectly — a warning sign, not an achievement, since a rule flexible enough to fit any past sample perfectly is usually too flexible to predict anything.
BiasWhat goes wrongThe fix
Survivorship biasDelisted, merged, or failed companies vanish from today's stock list, so the backtest sample only contains "winners" that survivedBuild your stock list from what was actually listed and trading on the test date, including names later merged, suspended, or delisted
Look-ahead biasA rule uses information (audited EPS, final AGM decisions, revised guidance) that was not yet public on the decision dateTimestamp every input; use only unaudited quarterly figures, disclosures, and prices genuinely available as of that date
OverfittingA rule's weights or thresholds are adjusted repeatedly until they perfectly fit one historical sampleFix the rule's logic before testing; if you must adjust it, retest only on data the adjustment has never seen (see Lesson 79.4)

Lesson 79.3 — Why NEPSE Specifically Is a Hard Market to Backtest

Even a backtester who avoids all three biases above still runs into a problem that is specific to NEPSE rather than to backtesting in general: there simply is not that much clean history to test against, and what history exists is not as independent, or as stable, as it looks.

Start with the raw amount of usable data. NEPSE traces its institutional history back to 1993, with its trading floor opening on 13 January 1994 — but for well over a decade after that, trading ran on an open-outcry floor system, with brokers calling out prices to one another rather than the market generating clean, timestamped electronic records. NEPSE moved to a semi-automated system only in 2007–08, and dematerialization of shares began in 2011 with the establishment of CDS and Clearing Limited. It was not until November 2017 that NEPSE rolled out its fully automated, broker-independent online trading system — the version of the market, with investors placing orders themselves through a Trading Management System, that most of today's participants would actually recognise. That means reliable, machine-readable, investor-verifiable daily price and volume data — the kind a backtest can actually trust down to the transaction — really only stretches back eight or nine years, not thirty.

REGULATORY DETAIL NEPSE's fully automated, broker-independent online trading system went live in November 2017; before that, the market ran on a semi-automated setup introduced in 2007–08, following decades of open-outcry floor trading — meaning genuinely clean, machine-verifiable daily data for most stocks realistically covers less than a decade.

Eight or nine years sounds workable until you notice the second problem: those years are not neutral, evenly-behaved history. They contain the tail end of the post-2016 boom (NEPSE's index hit an all-time high of 1,881.45 on 27 July 2016, before a bear market dragged it down toward roughly 1,100 by early 2019); a historic mania in 2020–2021 that pushed the index to a fresh all-time high of 3,111.09 on 3 August 2021, with single-day trading generating over 21 million shares changing hands and six scrips hitting the day's positive circuit breaker at once; and then a grinding, multi-leg crash through 2022 that took the index down through the 2,700, 2,600, and eventually the 2,000 mark, a fall of well over a year in length. Backtest across that whole window and you are really backtesting three or four completely different markets stitched end to end — a low-liquidity recovery, a euphoric bubble, and a prolonged unwind — and calling the average of all three "how the rule performs."

The third problem is more subtle and, in Keshav's experience, the one investors notice last: NEPSE does not actually give you as many independent tests as its stock count suggests. There are roughly 260 listed companies today, which sounds like 260 separate experiments. But a large share of them are commercial banks, development banks, and microfinance institutions whose share prices move together on the same handful of triggers — a Nepal Rastra Bank monetary policy announcement, an interest rate spread directive, a capital adequacy circular — and a large share of the rest are hydropower companies whose revenue, and therefore sentiment, swings with the same monsoon season and the same load-shedding or export-tariff news. If thirty bank stocks all rally or fall together because of one NRB circular, that is not thirty independent tests of your rule; it is closer to one test, repeated thirty times in slightly different costumes. Treating correlated stocks as independent data points is a quiet way of convincing yourself you have more evidence than you actually do.

WARNING Many NEPSE bank and hydropower counters move together on the same handful of triggers — an NRB policy circular, a monsoon season — so a backtest across thirty such stocks is closer to one real test repeated thirty times than to thirty independent tests.

Layer on top of that the structural rule changes NEPSE has gone through in this same short window — circuit breaker thresholds have been adjusted more than once, free-float and index-calculation methodology has changed, and the 2015 capital directive alone forced a wave of BFI mergers that reshaped which "stocks" even existed from one year to the next — and you get a market whose own rules of the game changed underneath the price history you are trying to test against. Keshav's blunt summary, scribbled in his ledger margin: "Eight years of data, four different markets, and the referee changed the rules twice." That is not a reason to abandon backtesting on NEPSE. It is a reason to do it with far more humility than a book on, say, the S&P 500's ninety years of data would require.

Lesson 79.4 — In-Sample vs Out-of-Sample Testing

The single discipline that would have caught Keshav's overfitting problem in Lesson 79.2 before it wasted a weekend is splitting the data in two and refusing to look at the second half until the first half is finished.

The two pieces of jargon here are simpler than they sound. In-sample data is the period you use to build or tune your rule — the period you are allowed to look at, argue with, and adjust your rule against as much as you like. Out-of-sample data is a separate period, one your rule has never seen and was never adjusted to fit, which you test the finished, frozen rule against exactly once. Think of it the way a schoolteacher thinks about practice exams versus the real board exam: you can revise your approach against as many practice papers as you like, but the actual board exam questions must be ones you have never seen in advance, or the exam proves nothing about whether you actually learned the subject.

Keshav restructured his test this way. He took his eight-and-a-half years of available NEPSE data (November 2017 through the recent close of Fiscal Year 2081/82) and split it: roughly the first five and a half years, through mid-2023, became his in-sample period, where he was allowed to build, adjust, and sanity-check the Canon Score's weights and his liquidity exit rule. The remaining period — from mid-2023 to the present — he sealed off entirely. He did not look at those prices while tuning anything. Only once his rule was completely fixed, weights and thresholds locked, did he run it forward against that untouched stretch, exactly once, and accept whatever came out.

KEY CONCEPT In-sample data is what you use to build and adjust a rule; out-of-sample data is a separate, untouched period you test the finished rule against exactly once — the discipline that catches overfitting before real money does.

The result humbled him usefully rather than painfully. His overfit, heavily-tuned version of the Canon Score — the one that had explained the in-sample years almost perfectly — performed distinctly worse out-of-sample than the simpler, less-tuned original version he had been using informally for the past year. The complicated version had been memorising the in-sample noise, not learning a real pattern, and so it had nothing useful to say about a period it had never seen. The simpler version, with fewer moving parts, held up closer to its in-sample performance. That gap between in-sample and out-of-sample results is itself a diagnostic: a rule whose out-of-sample performance collapses relative to its in-sample performance is telling you, plainly, that it was fitted rather than found.

This is also the single most common way retail "backtested" strategies fail once real money is on the line. An investor tunes a rule against the same data he later trades on — adjusting the entry threshold a little, adding a filter here, until the historical chart looks clean — and then is baffled when the "proven" rule underperforms going forward. It was never tested going forward. It was polished backward, against the only data it was ever allowed to see, and then unleashed on data that, from the rule's point of view, might as well be a different market entirely.

PRACTICAL TOOL Before tuning any rule against NEPSE history, physically set aside the most recent one to two years of data in a separate file you do not open until the rule is completely finished — a low-tech but effective way to force genuine out-of-sample discipline.

Lesson 79.5 — A Practical, Honest Backtesting Workflow

None of this requires software Keshav did not already own. He built his entire backtest in a spreadsheet, and the workflow he settled on — after his false starts — is one any patient Nepali retail investor can run by hand.

Step one is picking one rule and writing it down so precisely that two different people, given the same data, would reach the same decision. "Buy strong banking stocks" is not a rule; it cannot be tested because it cannot fail. "Buy a commercial bank scoring 70 or above on the Canon Score, using only data available as of the first trading day after each quarterly result is published, and hold until the score drops below 55 or twelve months pass, whichever comes first" is a rule. It has an entry trigger, a data cutoff, and an exit trigger, all specific enough that ambiguity is removed.

Step two is defining entry and exit with that same precision, including exactly which price you would have transacted at (the next day's opening price is usually the honest choice, since you could not have traded at a closing price you had not yet seen) and exactly which data vintage feeds the decision (quarterly unaudited figures, not the audited annual figure that arrives months later — the look-ahead trap from Lesson 79.2).

Step three is walking the rule forward year by year, strictly in date order, using only information that existed on each decision date. Keshav did this literally with a ruler and a printed price chart at first, covering each stock's future price movement with a sheet of paper so he could not see it while deciding what the rule would have done on a given date — a low-tech but effective discipline against the temptation to let hindsight creep in.

Step four, the one Keshav found most uncomfortable, is recording every trade the rule generates, including the ones that would have been embarrassing. He had, without quite admitting it to himself, been quietly skipping two trades in his early drafts — one, a hydropower stock the rule said to buy that then dropped 30 percent on a court case involving its power purchase agreement, and another, a finance company the rule held onto for eleven months while it drifted to a loss before finally triggering the exit. Leaving those two trades out of his tally moved his average return from mediocre to good. Putting them back in was the entire point of the exercise.

CAUTION The trades that feel embarrassing to include — the rule-generated buy that was followed by bad news, the exit that came too late — are exactly the trades that make a backtest honest; quietly excluding them turns the exercise back into a story.

Step five is computing a small number of honest summary numbers rather than a single flattering headline return. Keshav settled on four: hit rate (the percentage of trades that were profitable), average gain on winning trades versus average loss on losing trades (so a high hit rate built on tiny wins and rare but brutal losses does not disguise itself as a good rule), and maximum drawdown (the worst peak-to-trough decline the rule's running account value would have experienced, which tells you whether you could have actually stomached holding through the rule's worst stretch). A single overall return percentage can be true and still misleading — it can be produced by one lucky trade dominating forty unlucky ones. The four numbers together are much harder to fool.

Here is a simplified extract from the backtest log Keshav actually kept, covering a handful of trades from his Canon-Score-based bank and hydropower rule:

Entry dateStockCanon Score at entryEntry price (Rs)Exit dateExit price (Rs)Result
2019 MangsirBank A742852020 Ashadh340+19.3%
2019 FalgunHydro B714102019 Ashadh (following)295-28.0%
2020 KartikBank C681902021 Baisakh410+115.8%
2021 ShrawanFinance D705202022 Chaitra340-34.6%
2022 MangsirBank A762502023 Ashadh275+10.0%

Five trades is far too small a sample to draw conclusions from, and Keshav's real log ran to several dozen — but the format matters more than the count here: every row has a precise entry trigger, a precise exit trigger, and an outcome recorded whether it flatters the rule or not. That is the difference between a backtest log and a highlight reel.

PRACTICAL TOOL Keep a running backtest log with one row per trade — entry date, score or signal value at entry, entry price, exit date and price, and result — and update it before checking whether the trade was a winner, so the temptation to "forget" a bad one never gets the chance to operate.

Lesson 79.6 — The Limits of Backtesting

After several months of this work, Keshav reached a conclusion he found both satisfying and uncomfortable: the Canon Score, tested honestly against the out-of-sample period, performed reasonably — better than a naive buy-anything approach, with a hit rate around six in ten and a drawdown he judged tolerable — but nowhere near as spectacular as his first, biased attempt had suggested. He decided that was, in fact, the correct amount of confidence to have in it.

What a backtest can honestly tell you is narrow but real: that a rule was not obviously wrong across the specific historical stretch you tested, under the specific conditions that stretch happened to contain. It is evidence the rule is not pure fantasy. It is not, and can never be, proof the rule will keep working in conditions that stretch never contained. This matters enormously for NEPSE specifically, because the market's entire electronic-data history — the eight or nine years Keshav actually had to work with — has never yet contained a genuine, prolonged, multi-year bear market tested against today's participant base. Demat account holders grew from roughly 1.48 million in FY 2018/19 to nearly 3.79 million by FY 2020/21, and to close to 4.9 million since — meaning a large majority of the people currently trading NEPSE opened their accounts during or after the 2020–2021 mania and have only ever personally experienced the sharp-but-comparatively-brief 2022 correction that followed it. Nobody's backtest, however carefully built, can prove how a rule — or a market full of these investors — behaves in a slower, multi-year grind down, because that regime has not happened yet inside the clean data anyone can actually test.

CAUTION A backtest can show a rule was not obviously wrong in the past; it cannot prove the rule will survive a market regime — like a genuine multi-year bear market with today's much larger, much younger participant base — that has not yet occurred in NEPSE's clean electronic-data history.

This is why backtesting belongs in this book as a necessary discipline rather than a final answer. It replaces "it feels right" and "it worked for my friend" with something falsifiable, which is real progress. It cannot replace ongoing humility about the fact that markets, and NEPSE in particular, keep generating conditions nobody's historical sample has seen before. Keshav's own plan, once he finished this first honest pass, was not to bet his full retirement savings on the Canon Score with newfound certainty. It was to size his positions the way Part XII already taught him — respecting ADV-based liquidity limits regardless of how good the backtest looked — and to keep the backtest log running indefinitely, adding every new trade as it happens, so the rule keeps being tested against a market that keeps writing new history.

Chapter recap

This chapter opened Part XV by drawing a hard line between believing a rule works and actually having evidence that it did. Backtesting, in plain terms, means taking a precisely defined rule — a score, an entry signal, an exit trigger — and running it mechanically against historical data, using only the information that was genuinely available at each decision point, and then recording every result honestly. It is fundamentally different from "the strategy feels right" or "it worked for my friend," both of which are stories built after the fact rather than tests committed to before it.

We walked through the three biases that quietly wreck amateur backtests: survivorship bias, where delisted or merged companies vanish from today's stock lists and make the past look safer than it was; look-ahead bias, where a rule accidentally uses information — like full-year audited earnings — that would not actually have been available on the decision date, when only unaudited quarterly figures existed; and overfitting, where a rule is tuned and re-tuned until it perfectly explains one historical sample, at the cost of describing nothing repeatable at all. We then looked at why NEPSE specifically makes all of this harder than in older, larger markets: genuinely clean electronic data realistically covers less than a decade following the November 2017 rollout of fully automated trading; that short window already contains at least three distinct regimes — a post-2016 bear market, the 2020–2021 mania that peaked above 3,100 on the index, and the grinding 2022 correction; and a large share of NEPSE's roughly 260 listed companies move together on the same handful of triggers, so the market offers far fewer truly independent tests than its headline stock count suggests.

The chapter's central discipline was the split between in-sample and out-of-sample testing — building and tuning a rule on one period, then testing the frozen, unmodified rule exactly once on a separate period it has never seen. Skipping this step, and instead polishing a rule against the very data you later trade on, is the most common single reason retail "backtested" strategies disappoint once real money is involved. From there we built a practical, honest workflow any investor can run by hand or in a simple spreadsheet: define the rule with enough precision that two people would make the same call from the same data; walk it forward strictly in date order using only information available at each point; record every trade including the embarrassing ones; and judge the result using a small set of honest numbers — hit rate, average gain versus average loss, and maximum drawdown — rather than one flattering headline return.

Woven through all six lessons was Keshav Raj Pyakurel, a retired Nepal Rastra Bank supervision officer who spent a year using the Canon Score informally, then spent several more months testing it honestly against eight-plus years of NEPSE history before deciding how much of his retirement savings it deserved. His near-miss with overfitting, his discovery of the look-ahead trap hiding in audited EPS, and his discomfort at almost quietly dropping two losing trades from his log are not exaggerations for effect — they are the ordinary, specific ways a careful, numbers-literate investor can still fool himself, and the ordinary, specific disciplines that catch it.

The chapter closed on backtesting's real limit: it can tell you a rule was not obviously wrong across the history you tested, but it cannot prove the rule will hold up in a market regime that history has not yet produced — and NEPSE's clean electronic record, dominated by participants who joined during or after the 2020–2021 mania, has never yet contained a genuine, prolonged multi-year bear market. That gap is exactly where Chapter 80, "Calibrating Your Scoring Model Against NEPSE History," picks up the thread — taking the backtesting mindset built here and applying it directly to the Canon Score itself, dimension by dimension, to find out which of its seven components have actually earned their weight in NEPSE's real history and which have simply never yet been tested by conditions severe enough to matter.

Primary data sources Figures, rates and rules referenced in this chapter can be verified against the primary sources: Nepal Rastra Bank (monetary policy, credit and BFI data), SEBON (regulation and issue approvals), NEPSE (prices, indices and turnover), CDSC (settlement and demat data) and Inland Revenue Department (tax rates and rulings). If a figure here disagrees with the primary source, trust the primary source and tell me.