Part XV · Chapter 82

Measuring and Managing Model Drift

First published 25 Aug 2026 · Last verified 29 Aug 2026

Rukmini Thapa had not touched the Canon Score's weightings in fourteen months. That was, in itself, unusual for her. She had spent thirty-one years at Nepal Rastra Bank reviewing bank balance sheets, and the habit of a supervisor is to tinker — to tighten a ratio here, flag an exception there, never quite leave a rule alone. But when she retired and turned to building her own scoring system for NEPSE stocks, she had made herself a promise: calibrate it properly once, using the discipline she had learned two chapters ago, and then leave it alone unless she had real evidence, not a hunch, that it needed to change.

It was July 2026, and she was looking at a chart that made her deeply uncomfortable. The chart tracked something simple: every quarter for the last three years, she had ranked NEPSE stocks by their Canon Score and checked whether the top-scoring group actually went on to outperform the bottom-scoring group over the following ninety days. For most of that stretch, the gap between the two groups — what she'd started calling the "spread" — had hovered in a healthy, reassuring band. In the most recent two quarters, it had nearly vanished. Her score was still producing numbers. It had simply stopped telling her anything useful.

This chapter is about what Rukmini was looking at, and what she did next. It is the closing chapter of Part XV, and it exists because of an uncomfortable truth the last three chapters have been building toward: a properly backtested score and a properly calibrated set of exit rules are not permanent achievements. They are perishable. They decay. The market you calibrated against in 2021 is not the market you are trading in 2026, and a tool that does not get checked against that fact will quietly go stale while still looking, on the surface, exactly as confident as the day you built it.

Lesson 82.1 — What "Model Drift" Actually Means

Start with the plainest possible definition. Model drift is what happens when a rule that used to work well keeps running exactly as designed, while the world it was designed for changes underneath it — so the rule's outputs slowly stop matching reality, even though nothing about the rule itself has "broken."

Think of a tailor who takes your measurements once and cuts every shirt from that pattern for the next ten years. The pattern was not wrong the day it was made — it fit you perfectly then. But you do not stay the same size forever. Somewhere along the way, the shirts stop fitting, not because the tailor made a mistake, but because you changed and the pattern didn't. Nobody notices the exact day it stopped fitting; it happens gradually, one shirt at a time, until eventually every shirt in the closet feels slightly wrong and you can't say when that started.

A backtested investing rule is a pattern cut to fit a particular market at a particular time. The Canon Score, as calibrated in Chapter 80, was cut to fit NEPSE roughly as it behaved across 2018 to 2021 data — a market with a certain mix of banks, hydropower companies, microfinance institutions, and manufacturing firms; a certain size and temperament of retail investor base; a certain regulatory backdrop. The exit rules calibrated in Chapter 81 were cut to fit NEPSE's circuit-breaker mechanics and liquidity patterns as they existed when that testing was done. Neither of those things is fixed in stone. NEPSE in 2026 is not the same market, and pretending the pattern still fits because it fit once is exactly the kind of complacency this book has tried to train out of you since Part XV began.

KEY CONCEPT Model drift is not a sign your rule was wrong when you built it. It is a sign the market has moved and your rule hasn't moved with it. The two require completely different responses — one calls for humility about your original work, the other calls for an honest, unemotional update.

It helps to separate drift from two things it is often confused with. It is not the same as a single bad quarter, which can happen to a perfectly sound rule simply through ordinary variance — more on that distinction in Lesson 82.5. And it is not the same as the rule being poorly built in the first place, which is a calibration failure Chapter 80 already dealt with. Drift specifically describes decay over time in a rule that was sound when built. The tailor's pattern was cut correctly; you grew.

The uncomfortable part, and the reason this chapter exists at all, is that drift is invisible from the inside of any single decision. Rukmini's Canon Score, on any given day, still produced a number between 0 and 100 for any stock she typed in. It still felt authoritative. Nothing about using it that day would have told her the number meant less than it used to. Drift only becomes visible when you deliberately step back and measure performance over time — which is precisely the discipline the rest of this chapter builds.

Lesson 82.2 — Four Ways NEPSE Changes Under Your Feet

Drift does not happen for mysterious reasons. It happens because specific, identifiable things about a market change. On NEPSE, four sources of change matter most, and a Canon Score user should be able to name all four without hesitation.

The first is a regulatory shock — a rule change from Nepal Rastra Bank (NRB, the central bank) or the Securities Board of Nepal (SEBON, the capital markets regulator) that changes what "good" looks like almost overnight. The textbook example, and one worth knowing in real detail because it will recur in this book's case studies, is NRB's paid-up capital mandate for commercial banks. In 2015, NRB directed commercial banks to raise their minimum paid-up capital — the base equity cushion a bank must hold before it can call itself a "class A" commercial bank — roughly fourfold, from Rs 2 billion to Rs 8 billion, with a deadline of mid-2017 (the end of the Nepali fiscal year 2073/74). Banks that could not raise that capital organically were pushed into mergers, rights issues, and bonus-share issuances at a scale NEPSE had never seen. The number of listed commercial banks fell sharply as weaker institutions were absorbed into stronger ones, and every surviving bank's balance sheet looked structurally different on the other side of that deadline than it had before it.

REGULATORY DETAIL NRB's 2015 directive lifting commercial bank paid-up capital requirements from Rs 2 billion to Rs 8 billion, with compliance due by mid-2017, triggered one of the largest waves of bank mergers and forced capital-raising in NEPSE's history. Any scoring dimension built on "does this bank clear a capital threshold" measured something meaningfully different before and after that window.

Now imagine a capital-adequacy dimension in a scoring model built to reward banks that comfortably cleared a capital bar. Before the mandate, that dimension usefully separated the strong banks from the weak ones — the spread between them was wide. After every surviving bank was forced to clear the same much higher bar, nearly all of them cluster near the top of that dimension. The dimension hasn't become wrong; it has become uninformative, because the thing that used to distinguish banks from each other no longer distinguishes much of anything. This is exactly the kind of drift Rukmini eventually found, and we will return to it directly in Lesson 82.3.

The second source is sector composition shift — a change in what a "typical" NEPSE-listed company even looks like. Hydropower listings have expanded dramatically in recent years; by 2025, hydropower companies made up the largest single block of newly listed and pipeline companies on NEPSE, with dozens of hydropower IPOs approved by SEBON and a multi-billion-rupee pipeline of further hydropower issuances still to come. A scoring model built when banks and manufacturing firms dominated the index will have been implicitly calibrated on the financial patterns of those sectors — steady earnings, established capital structures, long trading histories. Hydropower companies behave differently: many are pre-revenue or early-revenue at listing, dependent on monsoon-driven river flows for output, carry different debt structures tied to project financing, and often list with thin early trading volumes. A model that quietly assumed "most stocks look like a bank" will misjudge a growing share of the market it is now being asked to score.

CASE IN POINT Hydropower listings and pipeline IPOs became the largest single sector by count of new NEPSE companies by 2025. A Canon Score built mostly against bank and manufacturing data from 2018-2021 is now being asked to judge a market with a meaningfully different sector mix than the one it learned from.

The third source is participant base change — who is actually doing the buying and selling. NEPSE's investor base has grown enormously; demat accounts (the electronic share-holding accounts required to trade) numbered in the low millions a few years ago and had climbed toward roughly eight million by 2026, swelled by remittance-fed household savings and a post-pandemic wave of first-time retail participation. A market with a much larger base of newer, smaller, more sentiment-driven retail traders can behave differently at the margin than one dominated by a smaller pool of more experienced participants — prices can move faster on rumour, IPO listings can pop and fade more sharply, and a scoring dimension that weights "market sentiment" or "momentum" was calibrated against the crowd psychology of a smaller, different crowd.

The fourth source is the plainest: calendar time itself. A rule calibrated on 2018-2021 data is, by definition, being asked to perform in a market it has never seen — different government, different remittance flows, different global interest-rate backdrop, different investor mood. Even with no single dramatic regulatory shock or sector shift, enough ordinary years passing is itself a source of drift, the way a ten-year-old road map is out of date even if no earthquake ever hit the city.

WARNING None of these four sources of drift show up as a single dramatic event you can circle on a calendar and say "the model broke here." They accumulate quietly. The only way to catch them is to measure performance on a schedule, not to wait until you feel that something is wrong.

Lesson 82.3 — Catching Drift Before It Costs You

The instinct many investors have is to wait for pain — to notice drift only after a string of losses forces the question. That is the most expensive way to find out. By the time a decayed model has cost you real money, it has usually been decayed for a while. The alternative is to track a small number of honest, boring scorecards on a fixed schedule, so that decay shows up as a number moving in the wrong direction long before it shows up as a loss.

Three scorecards do most of the work.

The first is quintile spread. Take every quarter's Canon Score rankings, split the scored universe into five equal groups (quintiles) from highest score to lowest, and track the actual subsequent-quarter return of the top quintile against the bottom quintile. A healthy score produces a meaningful, persistent gap — the top quintile should meaningfully outperform the bottom quintile more quarters than not. When that gap narrows quarter after quarter, the score is losing its ability to discriminate between good and bad stocks, which is the entire point of having a score.

The second is hit rate — the percentage of stocks the score rated in its top band that actually delivered above-median returns over the following quarter. A score with a stable, useful hit rate somewhere comfortably above chance is doing its job. A hit rate sliding steadily downward over several consecutive quarters, even without a single catastrophic quarter, is the quiet signature of drift.

The third, specific to the exit-rule work of Chapter 81, is trigger frequency relative to realised volatility — how often the stop-loss or exit rule fires relative to how much the underlying stock is actually moving. If exit rules are getting triggered by ordinary day-to-day noise far more often than they used to relative to that noise, either the market's volatility character has shifted (plausible, given the retail base change discussed above) or the rule's thresholds no longer match current conditions.

Table 82.1 shows the kind of scorecard Rukmini had actually been keeping, quarter by quarter, for three years — the exact data that let her catch her drift problem before it cost her a bad year rather than after.

QuarterTop-quintile avg returnBottom-quintile avg returnSpreadScore hit rateExit-rule trigger rate
2023 Q311.2%2.1%9.1 pts68%1 per 14 trading days
2024 Q110.4%1.8%8.6 pts65%1 per 13 trading days
2024 Q39.8%3.0%6.8 pts63%1 per 12 trading days
2025 Q18.1%4.4%3.7 pts58%1 per 9 trading days
2025 Q36.9%5.2%1.7 pts54%1 per 7 trading days
2026 Q16.2%5.6%0.6 pts51%1 per 6 trading days

Read across that table the way Rukmini did. The spread between best-rated and worst-rated stocks did not collapse in one bad quarter — it eroded steadily across nearly three years, from a healthy 9 points down to a statistically meaningless 0.6 points. The hit rate slid from a comfortably above-chance 68 percent toward a coin-flip 51 percent. And the exit rule, calibrated in Chapter 81 against 2018-2021 volatility patterns, was firing roughly twice as often relative to trading days by early 2026 as it had three years earlier — a sign that ordinary price movement itself had gotten choppier, plausibly tied to the much larger and more sentiment-driven retail base discussed in Lesson 82.2.

PRACTICAL TOOL Keep exactly three numbers on a running spreadsheet, updated every quarter: quintile spread, hit rate, and exit-rule trigger rate relative to trading days. Three honest numbers, tracked consistently, will tell you more about drift than any amount of "the market feels different lately" intuition.

None of these three numbers requires sophisticated statistics or expensive data. They require the same discipline Chapter 79 first introduced: writing down what you expected in advance, and then checking the actual result against it, quarter after quarter, without skipping the quarters where the answer is inconvenient.

Lesson 82.4 — The 90-Day Review Cadence

Measuring is only half the job. The other half is having a fixed process for acting on what you measure — one that runs on a schedule rather than on emotion. This book's broader framework calls for a 90-day feedback loop, and Part XV's closing lesson on process is simply this: apply that same cadence to the Canon Score and the exit rules themselves, not only to your individual trades.

A proper 90-day model review has four steps, done in order, every quarter, whether or not anything feels wrong.

First, update the three scorecards from Lesson 82.3 with the latest quarter's data. This step is mechanical and should take under an hour if you've kept the spreadsheet current. Resist the temptation to skip it in a quarter where you suspect the numbers will look bad — those are exactly the quarters the review exists for.

Second, re-run a lightweight backtest on the most recent quarter using only information that was actually available at the time — a point-in-time backtest, meaning you score stocks using only the data an investor genuinely had in hand on that day, not data that arrived later (an earnings report published after your hypothetical decision date, for instance). This guards against the same look-ahead bias Chapter 79 warned about in the original calibration work; a review that quietly uses future information to judge past decisions will always look better than it should.

Third, ask a specific, narrow question about each of the four drift sources from Lesson 82.2: has any regulatory change occurred this quarter that alters what a scoring dimension is actually measuring? Has the sector mix of newly listed or actively traded stocks shifted meaningfully? Has anything observable changed about the participant base (a surge in new demat accounts, a notable change in average holding period, a wave of IPO-driven speculative activity)? And, simply, how much calendar time has passed since the last full recalibration? Answering these four questions in writing, every quarter, turns a vague feeling of unease into a specific, checkable claim.

Fourth, and only after the first three steps are done, decide explicitly whether any weight or threshold needs adjustment — and write down the decision and the reasoning, even if the decision is "no change this quarter." That written record matters enormously, because it is what lets you tell, a year later, whether you have been making disciplined adjustments or slowly overfitting one quarter at a time.

KEY CONCEPT A 90-day model review is not a chance to redesign your system. It is a chance to check three numbers, ask four questions, and make a small number of narrow decisions — then close the file until the next quarter.

Rukmini's own review, the one that opened this chapter, followed exactly this process. Her scorecards had already flagged the declining spread. Her point-in-time backtest of the most recent quarter confirmed the capital-adequacy dimension of her score was assigning nearly every commercial bank a score between 82 and 96 — a range so narrow it was barely distinguishing anything. Working through the four drift-source questions, the regulatory answer jumped out immediately: the Rs 8 billion paid-up capital mandate, a full decade old by 2026, had done its work so thoroughly that virtually every surviving commercial bank now cleared it comfortably. A dimension designed to reward banks for clearing a bar that used to separate the strong from the weak was now rewarding everyone roughly equally, because the bar itself had become the industry floor rather than a meaningful hurdle.

Lesson 82.5 — Signal vs Noise: Not Every Bad Quarter Is Drift

Here is where the discipline gets genuinely difficult, and where this chapter connects directly back to the overfitting warning Chapter 80 spent an entire lesson on. A single disappointing quarter is not, by itself, evidence of drift. Markets have ordinary bad stretches for reasons no reasonable scoring model was ever built to predict — a one-off political disruption, a monsoon shock that hits hydropower generation unevenly, a temporary liquidity squeeze around a major festival period, a single large institutional seller distorting one sector's prices for a few weeks. Treating every such episode as proof the model is broken, and re-tuning weights in response, is not model maintenance. It is overfitting in slow motion — the exact trap Chapter 80 warned against, just spread across quarters instead of concentrated in one backtest.

The practical challenge is telling the two apart in real time, and it comes down to persistence and mechanism. Genuine structural drift shows up as a trend across several consecutive quarters and can usually be traced to an identifiable, permanent change in what a dimension measures — a regulatory bar that has become universal, a sector shift that has permanently changed the composition of what gets scored. Ordinary noise shows up as a single bad quarter with an identifiable one-off cause, followed by a return to the prior pattern once that cause passes.

Signal (genuine drift)Noise (ordinary variance)
Trend persists three or more consecutive quartersConfined to a single quarter, then reverts
Traceable to a specific, permanent structural change (regulatory mandate, sector shift, participant base shift)Traceable to a one-off event (political disruption, weather shock, single large seller) with no lasting mechanism
Affects a specific dimension in a way consistent with the identified cause (capital-adequacy scores compress after a capital mandate saturates the industry)Affects returns broadly and indiscriminately, without a clean link to any one scoring dimension
Point-in-time backtest of the recent quarter confirms the same weakness using only period-appropriate dataPoint-in-time backtest shows the dimension behaved normally; the return shortfall came from something outside what the dimension was ever built to measure
Other independent evidence corroborates it (news of the regulatory change, visible sector composition data, demat account statistics)No external corroborating structural change can be found
CAUTION If your response to a single rough quarter is to immediately adjust a weight or threshold, ask yourself honestly whether you are managing drift or simply overfitting to the most recent three months of noise. The correct response to an unexplained bad quarter, absent a persistent trend or identifiable structural cause, is usually no change at all — just another data point banked for the next review.

This is precisely why Rukmini did not touch her score the first time the spread narrowed slightly, back in early 2025. One soft quarter told her nothing. It was the trend across six consecutive quarters, combined with a specific and traceable regulatory mechanism a decade in the making, that gave her confidence she was looking at genuine drift rather than noise. She also checked the mechanism specifically against her exit-rule trigger data, and found a second, independent piece of corroborating evidence: the exit rule firing roughly twice as often relative to trading days lined up with the demat-account growth and retail participation surge discussed in Lesson 82.2 — not with any single news event she could point to for one bad month.

Her actual adjustment, once she was confident it was warranted, was narrow. She did not rebuild the Canon Score from scratch. She replaced the capital-adequacy dimension's binary "does this bank clear Rs 8 billion" threshold — now nearly meaningless since almost every surviving bank clears it — with a relative measure comparing each bank's capital adequacy ratio (a bank's capital as a percentage of its risk-weighted assets, a standard regulatory soundness measure) against its peers today, restoring some spread to a dimension that had flattened out. She left every other dimension of the score untouched, and she made a small, explicit note to revisit the sentiment-weighting dimension's calibration at her next scheduled review rather than acting on a single quarter's signal from the retail-base data.

Lesson 82.6 — Conviction and Humility: Closing Part XV

Four chapters ago, Part XV opened with a promise: that a stock-picking system built on hope and gut feeling is worse than useless, and that discipline — testing, calibrating, checking your work against reality rather than your memory of reality — is what separates a durable investing process from a lucky streak that eventually runs out. Chapter 79 taught you how to test a rule honestly. Chapter 80 applied that testing to the Canon Score itself. Chapter 81 applied it to the exit rules that protect you on the way out. This chapter closes the loop: even a properly tested rule is not a finished object. It is a living tool, and living tools need periodic, honest re-examination or they quietly decay into shapes that no longer fit the market wearing them.

The discipline this chapter asks of you is really the ability to hold two things at once, in some tension with each other, without letting either one win outright. The first is conviction: enough confidence in a properly calibrated Canon Score and a properly tested exit rule to actually use them consistently, quarter after quarter, without abandoning them the moment a single result disappoints you. A tool you redesign every time it has a bad month is not a tool at all — it is an excuse to keep chasing whatever happened most recently, which Chapter 80 already showed you is just overfitting wearing a different hat. The second is humility: enough honesty to keep measuring, on a fixed schedule, whether the tool still deserves that confidence — and enough willingness to make a narrow, well-evidenced adjustment when the scorecards, not your gut, say the ground has genuinely shifted.

Rukmini's year illustrates both halves. She did not touch her score for fourteen months despite plenty of ordinary quarterly noise along the way — that was conviction, and it kept her from overfitting to every rough patch. But when six consecutive quarters of narrowing spread lined up with a specific, traceable, decade-old regulatory mandate finally saturating an entire industry, she did not simply keep the faith and hope the pattern reversed on its own — that would have been complacency dressed up as discipline. She measured, she distinguished signal from noise using the table in Lesson 82.5, and she made one narrow, well-justified change to one dimension of her score, leaving the rest alone. That is what a mature relationship with a calibrated model looks like: not blind loyalty, and not constant tinkering, but a fixed process for occasionally, deliberately, earning the right to change your mind.

WARNING A model you never revisit and a model you revisit every week fail for the same underlying reason: neither is actually being checked against evidence on a disciplined schedule. Drift management lives in the middle, at a fixed cadence, with a written record of what you found and why you acted or didn't.

Chapter recap

This chapter defined model drift in plain terms: a scoring or exit rule that worked well when it was calibrated can quietly stop working as the market it was built against changes shape, not because the original work was flawed but because the world it measured no longer exists in the same form. On NEPSE specifically, four forces drive that change — regulatory shocks like NRB's decade-old paid-up capital mandate that pushed commercial banks from a Rs 2 billion to an Rs 8 billion minimum and, in doing so, eventually flattened a capital-adequacy scoring dimension that used to usefully separate strong banks from weak ones; sector composition shifts, most visibly the surge of hydropower IPOs that has made hydropower the largest single block of new NEPSE listings and changed what a "typical" scored company even looks like; participant base changes, as NEPSE's demat account base swelled toward roughly eight million accounts and reshaped how sentiment-driven price action can be at the margin; and the simple passage of calendar time, which alone guarantees that a rule calibrated on 2018-2021 data faces conditions it never saw.

The chapter then laid out how to detect that decay without waiting for a painful loss to reveal it: three honest, boring scorecards — quintile spread between top- and bottom-rated stocks, the score's hit rate, and how often exit rules trigger relative to ordinary volatility — tracked every quarter on a fixed schedule, exactly like the one shown in Table 82.1, where a slow decade-long erosion from a 9-point spread to a 0.6-point spread told the real story long before any single quarter's loss would have. It paired that measurement habit with a concrete four-step 90-day review process: update the scorecards, re-run a point-in-time backtest of the latest quarter using only period-appropriate information, ask the four drift-source questions directly, and only then decide explicitly whether to adjust anything — recording that decision even when it is "no change."

Just as important as knowing how to detect drift is knowing how not to overreact to it. The chapter drew a sharp line between genuine structural drift — persistent across several quarters, traceable to an identifiable and permanent change like a saturated regulatory threshold — and ordinary variance, a single rough quarter caused by a one-off event the model was never built to predict. Treating every disappointing quarter as proof of failure and re-tuning weights in response is not vigilance; it is the same overfitting trap Chapter 80 warned against, merely spread across calendar time instead of concentrated inside one backtest.

The narrative thread followed Rukmini Thapa, the NRB-veteran investor who calibrated her Canon Score in Chapter 80, through her first full year of disciplined quarterly reviews — showing her catching a real, decade-in-the-making case of drift in her capital-adequacy dimension precisely because she had been tracking the right three numbers on a fixed schedule, and showing her making one narrow, well-evidenced fix rather than a wholesale rebuild.

That balance of conviction and humility is this chapter's, and Part XV's, final lesson: a properly built model is never "done." It earns continued use only through periodic, honest re-testing, and the investor's real job is holding both a working system in her hands and a healthy suspicion of it in the back of her mind at the same time.

With that, Part XV — Calibration and Backtesting — comes to a close. The book now turns from the machinery of building and testing models to the work of actually using them. Part XVI, Mastery, opens with eight full worked case studies that take everything built across Parts XIII through XV — the seven-dimension Canon Score, its calibration discipline, its exit rules, and now its drift-management process — and apply all of it, start to finish, to real company archetypes an investor will actually encounter on NEPSE. Chapter 83, Case Study 1 — A Commercial Bank, is the first of those eight studies, and it will put the very capital-adequacy dimension this chapter just revisited under the microscope: scoring a real commercial bank archetype end to end, showing exactly how a decade of NRB capital mandates, the merger wave that followed, and the drift-adjusted scoring approach from this chapter all come together in a single live decision.

Primary data sources Figures, rates and rules referenced in this chapter can be verified against the primary sources: Nepal Rastra Bank (monetary policy, credit and BFI data), SEBON (regulation and issue approvals), NEPSE (prices, indices and turnover), CDSC (settlement and demat data) and Inland Revenue Department (tax rates and rulings). If a figure here disagrees with the primary source, trust the primary source and tell me.