Part XV · Chapter 80

Calibrating Your Scoring Model Against NEPSE History

First published 25 Aug 2026 · Last verified 29 Aug 2026

Bimal Sharma spent thirty-one years at Nepal Rastra Bank, most of them in bank supervision, reading balance sheets the way other people read newspapers. When he retired in 2019, he brought that habit with him into the stock market, and it was this habit that had, in Chapter 79, led him to distrust his own excitement about the Canon Score before he trusted it with real money. He had tested the general discipline of backtesting there — the traps of survivorship bias, look-ahead bias, and overfitting, and the difference between checking an idea on data you have already seen versus data you have not. Now he turned that same discipline toward a narrower and, in some ways, harder question: not "does backtesting work," but "does my score itself add up correctly." The Canon Score he had been using — the 100-point, seven-dimension system built in Chapter 64, weighing Financial Strength & Profitability at 20 points, Governance & Promoter Behaviour, Valuation Reasonableness, Sector & Business Model Durability, and Growth Trajectory at 15 points each, and Liquidity & Tradability and Dividend & Capital Return Discipline at 10 points each — had felt right to him for two years. What he had never done was go back and check, sector by sector, whether he was actually applying Chapter 65's sector-specific adjustments with real discipline every time, or whether he had quietly been scoring every company — bank, hydropower plant, or microfinance institution alike — against the same generic bands underneath, while telling himself he was "keeping the sector context in mind." This chapter follows Bimal as he checks that, dimension by dimension, using real NEPSE history rather than gut feeling, and finds that at least one sector's adjustment had been quietly skipped for two years running.

Lesson 80.1 — What Calibration Really Means

Backtesting a trading rule, as Chapter 79 covered, asks a yes-or-no question about behaviour: if I had bought when the rule said buy and sold when it said sell, would I have made money, and would that have held up outside the exact period I tested it on? Calibrating a score is a different, quieter kind of question. It does not ask whether the Canon Score, taken as a whole, would have picked winners. It asks whether the internal architecture of the score — the number of points assigned to each of its seven dimensions, and how faithfully each dimension is actually being scored for the sector in front of you — reflects what has actually mattered for NEPSE outcomes, as opposed to what sounds like it should matter.

Think of the Canon Score as a recipe rather than a single dish. A recipe can produce a dish that tastes fine even when one ingredient is present in the wrong quantity, because the other ingredients compensate. Calibration is the process of tasting each ingredient on its own, not just the finished plate. Chapter 64 assigns 20 points to Financial Strength & Profitability and 15 points to Sector & Business Model Durability — an implicit claim that raw profitability and balance-sheet safety matter somewhat more than durability and moat in separating a good long-term NEPSE holding from a bad one. But there is a second, quieter claim buried inside that first one, and it is the one Bimal had never actually tested: Chapter 65 says that for a hydropower company, a microfinance institution, or a bank, several of those seven dimensions need a genuinely different ruler — different sub-ratios, different thresholds, sometimes an entirely different question. Was Bimal actually reaching for that different ruler every time, or was he scoring every company against the same generic bands and calling it "sector awareness" in his head? That claim had never been tested against Nepal's own trading floor. Calibration is the act of testing it.

KEY CONCEPT Calibrating a score means testing whether its point-weights and its sector-specific applications match what actually mattered for real outcomes — it is different from backtesting a trading rule, which tests only the buy/sell decision the score produces, not the score's internal architecture or how faithfully that architecture was actually applied.

This distinction matters because a score can pass a crude backtest — "stocks that scored above 70 mostly went up" — while still being badly calibrated underneath. A dimension could be scored using the wrong ruler for a whole sector, and the error could still wash out across a large enough sample, so the total looks fine even though the underlying reasoning was wrong every single time it was applied to that sector. A trader who only checks the final score against outcomes, the way Chapter 79 checked a trading rule, will never see this. Bimal, with his supervisory instincts, wanted to open the recipe up and check each ingredient. He started by writing out, on a single sheet of paper, what each of the seven Canon Score dimensions was supposed to be measuring, and — critically — which of them Chapter 65 says need a sector-specific rework rather than the generic Chapter 64 version. Financial Strength & Profitability (20 points) needed the heaviest rework of all: CAR, NPL ratio, and NIM for a bank; construction discipline pre-COD and season-adjusted margins post-COD for a hydropower company; provisioning coverage weighted above generic leverage for a microfinance institution. Governance & Promoter Behaviour (15 points) needed a heavy rework specifically for microfinance, where loan recycling and multiple-borrowing risk live outside the generic promoter-pledging checklist entirely. The other five dimensions needed lighter, more occasional adjustment.

Bimal's first honest realisation was that he had never once, in two years of using the score, gone back to check whether he was actually pulling out Chapter 65's specific sector guidance every time he scored a bank, a hydropower company, or a microfinance institution — or whether he had been scoring all three off the same generic Chapter 64 bands and simply trusting his own judgment to "adjust mentally" for sector, an adjustment that left no trail and that he had never once written down or tested against what actually happened to real NEPSE companies.

Lesson 80.2 — The Small-Sample Problem You Cannot Escape

Before Bimal could test anything, he needed to be honest about how much data he actually had to test with. NEPSE, as of the most recent sectoral count, lists 271 companies. On the surface, 271 data points sounds like a workable sample — enough to run some rough statistics on. But Bimal, from his NRB years, knew to ask a harder question: how many of those 271 are genuinely independent observations, and how many are just the same story told 91 times?

The sectoral breakdown makes the problem concrete. Of NEPSE's 271 listed companies, roughly 91 — about a third of the entire exchange — are hydropower companies. Microfinance institutions add another 50, roughly a fifth of all listings. Commercial banks number about 19, development banks about 16, finance companies about 20, and life and non-life insurers together add roughly 27. Manufacturing and processing companies, hotels, trading houses, and investment companies fill out the rest in much smaller numbers — manufacturing at around 22, hotels at only 7.

REGULATORY DETAIL NEPSE's roughly 271 listed companies break down, by rough sector share, into about 91 hydropower companies (a third of the exchange), 50 microfinance institutions (a fifth), and a combined 55 or so banks and development banks — meaning three sector clusters account for well over half of everything listed.

The trouble is that companies within a sector on NEPSE do not move independently of one another. Nearly all 91 hydropower companies generate electricity from rivers whose flow depends on the same monsoon, the same winter dry season, and the same glacial melt patterns; nearly all of them sell power under similar Power Purchase Agreements, or PPAs (long-term contracts that lock in the price a hydropower company is paid for its electricity, usually with the state utility), whose tariff structures were shaped by the same handful of regulatory decisions. When a dry winter reduces river flow across the country, it does not hit one hydropower company — it hits most of them at once, in the same direction, for the same reason. Similarly, nearly all commercial banks respond to the same interest-rate and capital-adequacy decisions from Nepal Rastra Bank (NRB, the central bank); nearly all microfinance institutions were exposed to the same rural-lending stress that built up in the early 2020s and worsened through 2025 and 2026. A "backtest" that scores all 271 NEPSE companies and checks outcomes is not really testing 271 independent scenarios. It is closer to testing three or four independent scenarios — one hydropower monsoon cycle, one banking-sector rate cycle, one microfinance credit cycle, one small and heterogeneous "everything else" cluster — each of which happens to contain many correlated repetitions of the same underlying event.

This is a subtler cousin of a problem statisticians call clustering, or lack of independence between observations. It means the effective sample size behind any Canon Score calibration exercise on NEPSE is not 271. It might realistically be closer to a dozen truly distinct historical episodes, once correlated companies are collapsed into the single event they are all reacting to.

WARNING Counting all 271 NEPSE-listed companies as 271 independent test cases is a statistical illusion — because roughly a third are hydropower firms reacting to the same monsoon and PPA cycle, and a fifth are microfinance institutions that were exposed to the same rural credit stress, the real number of independent scenarios behind any NEPSE-wide calibration is closer to a handful than to hundreds.

Bimal's conclusion from this was not that calibration was pointless — it was that any conclusion drawn from it had to be held loosely, and stated with appropriate humility. If he found that Chapter 65's microfinance-specific Governance and Financial Strength questions had correctly separated strong from weak microfinance institutions during the 2021-2026 stress period, that was one genuine data point about one genuine historical episode, not proof that his newly-added process would catch every future sector shock Nepal's capital market might produce. He wrote this caveat at the top of his calibration notes in block letters, precisely because he knew that after a few hours of satisfying pattern-matching, it would be tempting to forget it: one good match on one sector cluster is a clue, not a law.

Lesson 80.3 — Three Real Situations, Scored Blind

With that caution in place, Bimal built a practical method. He selected three real, verifiable NEPSE situations from the past decade — one from banking, one from microfinance, one from hydropower — chosen specifically because their outcomes were already public and could not be argued with. For each, he tried to reconstruct, as honestly as he could, what the Canon Score would have said before the outcome was known, using only information that would have been available at the time. Then he compared that pre-outcome score against what actually happened. This is the same in-sample discipline Chapter 79 described — except here the "rule" being tested was not a buy signal but the internal weighting of the score itself.

The first case was the 2013 merger that created NIC Asia Bank, formed when Nepal Industrial and Commercial Bank combined with Bank of Asia — the first-ever merger between two commercial banks in Nepali banking history. At the time of the merger, an investor scoring the combined entity would have been weighing genuinely uncertain integration risk against two banks with decent underlying fundamentals and management teams with a credible track record. Reconstructing the pre-merger picture using Chapter 65's bank-specific reading of Financial Strength (CAR and NIM in place of generic ratios) and Governance (related-party lending and loan concentration in place of a generic checklist), Bimal estimated the Canon Score would have landed in the low-to-mid seventies — solid marks on Financial Strength & Profitability and Governance & Promoter Behaviour for both underlying banks, a modest penalty on Sector & Business Model Durability for the genuine, temporary integration uncertainty a first-of-its-kind bank merger carried, and unremarkable marks elsewhere. The actual outcome: the combined bank went on to become Nepal's largest by customer base and balance sheet size, was recognised internationally, and sustained years of profitable growth. On this case, the score's direction was right, and specifically because Bimal had actually reached for Chapter 65's bank-specific rulers rather than scoring both banks on generic terms.

The second and third cases came from the microfinance sector, and this is where Bimal's confidence started to wobble. Nepal's microfinance institutions (MFIs — regulated lenders that provide small, mostly rural loans, often without traditional collateral) went through a well-documented stress period beginning around 2021-22 and worsening substantially by 2025-26, as rural loan demand cooled, over-indebtedness among borrowers became visible, and the sector's average non-performing loan ratio — the NPL ratio, meaning the share of loans not being repaid on schedule — climbed sharply, reaching roughly 11.35 percent sector-wide in one recent quarter, up from under 7 percent a year earlier. Eighteen microfinance companies crossed the 10 percent NPL threshold, with the worst performers — institutions in the Infinity and Dhaulagiri clusters, among others — reporting NPL ratios above 20 percent, in some cases approaching a quarter of their entire loan book. At the other end of the same sector, Chhimek Microfinance held its NPL ratio at roughly 2.3 percent throughout the same period, the lowest in the sector by a wide margin.

CASE IN POINT During Nepal's 2021-2026 microfinance stress episode, the weakest institutions reported non-performing loan ratios above 20 percent while Chhimek Microfinance held its ratio near 2.3 percent — the same sector, the same rural credit downturn, and a roughly tenfold difference in outcome.

Bimal tried to reconstruct pre-stress Canon Scores for a representative strong MFI (using Chhimek's public disclosures as a stand-in) and a representative weak one, using only what would have been visible before 2021 — loan book growth rates, published capital ratios, branch expansion pace, and management commentary in annual reports. The uncomfortable finding was that the two pre-stress scores came out close to each other, both somewhere in the mid-to-high sixties. Both institutions showed reasonable Financial Strength & Profitability on paper — decent growth, positive earnings, expanding branch networks — and neither showed obvious red flags on Governance & Promoter Behaviour, scored the way Bimal had always scored it: promoter shareholding stability, related-party transactions, disclosure timeliness. But that was exactly the problem. Chapter 65 is explicit that a microfinance institution's Governance dimension needs to be supplemented with two sector-specific questions the generic checklist never asks — whether the institution is recycling loans to disguise a rising NPL ratio, and whether its borrowers show signs of overlapping debt with other lenders in the same geography — and Bimal had never actually asked either question of either institution before 2021. He had also never applied Chapter 65's instruction to weight loan-loss provisioning coverage more heavily than generic leverage within Financial Strength & Profitability for a lender whose loan book is largely unsecured. The specific portfolio-quality signals that later proved decisive — loan concentration in overlapping rural districts, aggressive growth in loan officer headcount relative to oversight capacity, early upticks in loan rescheduling — were sitting exactly where Chapter 65 said to look for them. Bimal simply had not been looking there. The score, as he had actually been applying it, would not have told him which of these two institutions was headed for an NPL crisis and which was headed for the sector's best asset quality. That is a real calibration failure, not a hypothetical one, and Bimal wrote it down exactly that way rather than explaining it away.

The table below summarises all three test cases as Bimal recorded them.

SituationPre-outcome Canon Score (approx.)Dimension driving the scoreActual outcomeScore's verdict, in hindsight
NIC Asia Bank, 2013 merger (NIC Bank + Bank of Asia)Low-to-mid 70sGovernance and Financial Strength scored well using Chapter 65's bank-specific rulers; Durability lightly penalised for merger uncertaintyBecame Nepal's largest bank by customers and balance sheet; sustained profitable growth for a decadeCorrect call, for the right reasons
Chhimek Microfinance (pre-2021, stand-in for a well-run MFI)Mid-to-high 60sFinancial Strength (growth, earnings) scored well on generic terms; Chapter 65's loan-recycling, multiple-borrowing, and provisioning-coverage checks never actually appliedNPL ratio held near 2.3 percent through the 2021-2026 stress period, best in sectorRight outcome, but the score could not explain why in advance
Representative weak MFI, high-NPL cluster (pre-2021)Mid-60s, similar to ChhimekSame generic scoring as above — Chapter 65's MFI-specific questions never asked here eitherNPL ratio rose above 20 percent by 2025-26, among the weakest in the sectorWrong outcome — score failed to separate two institutions that turned out very differently
PRACTICAL TOOL To test a scoring model's calibration without needing a spreadsheet full of statistics, pick three to five real, already-resolved situations from different NEPSE sectors, reconstruct the score using only information available before the outcome, and compare against what actually happened — writing down honestly where the score would have been wrong, not just where it would have been right.

The fourth situation Bimal examined came from hydropower, and it exposed a different kind of miscalibration — not a dimension too weak to catch real danger, but a dimension too loud, reacting to the calendar rather than to anything company-specific. He picked a representative, already-operating (post-COD, in Chapter 65's terms) run-of-river hydropower company during a dry-season quarter, when reduced river flow cut generation sharply for several months, and scored its Financial Strength & Profitability the way he always had: off the single most recent reported quarter. That single quarter showed weak revenue, thin margins, and a return on equity that looked, in isolation, like a company in real trouble. Because nearly every run-of-river hydropower company on NEPSE was reporting a similarly weak dry-season quarter at the same time, the Canon Score would have marked down almost the entire sector at once — exactly the correlated, sector-wide reaction Lesson 80.2 warned about — even though the company's underlying PPA tariff was fixed by long-term contract and unaffected by the dry season, and nothing about its competitive position had changed at all. Chapter 65 says explicitly that a hydropower company's post-COD numbers must be read against the same quarter a year earlier, or on a trailing-twelve-month basis, precisely to strip out this wet-dry seasonality — and Bimal had never actually built that comparison into his own process. Once the following monsoon arrived on a normal schedule, generation and earnings recovered, and the single-quarter score penalty proved to have measured nothing durable — only which month it happened to be. This is the mirror image of the microfinance case: not a dimension whose generic version was too thin to catch a real problem, but a dimension being scored on the wrong window of time, for a sector where Chapter 65 had already said, in writing, which window to use instead.

Lesson 80.4 — Common Calibration Failure Patterns

Looking across all four situations together, Bimal could name three distinct failure patterns, and it was useful to him to give each one a name so he would recognise it faster the next time.

The first pattern is a dimension scored on the wrong window of time or the wrong ruler, so that it sounds important in theory but barely discriminates in practice. The hydropower case is the clearest example from Bimal's own test. Financial Strength & Profitability is intuitively the right dimension to carry real weight — of course whether a company is making safe, sustainable money matters — but scored off a single raw quarter for a sector like hydropower, where nearly every company's generation swings together on the same wet-dry cycle, the dimension does not help separate a good company from a bad one; it mostly just measures which month it is. A dimension that moves in lockstep across almost every company in a sector, simply because it was measured on the wrong time window, is not adding information a careful investor did not already have; it is adding noise dressed up as signal, even though the dimension itself — and its 20-point weight — is exactly the right one, once measured correctly.

WARNING A scoring dimension that swings the same way, at the same time, across nearly every company in a sector is not measuring company-specific quality — it is measuring the sector's shared weather. Before concluding a dimension is badly weighted, check whether it is simply being measured on the wrong time window for that sector.

The second pattern is the reverse: a dimension whose generic Chapter 64 version is scored correctly on its own terms, but which Chapter 65 says needs specific sector-specific questions added that never get asked in practice. Governance & Promoter Behaviour's generic checklist — promoter pledging, related-party transactions, disclosure timeliness — is a perfectly reasonable general-purpose measure. But in the microfinance case, the specific danger signals that mattered — loan recycling, multiple-borrowing and overlap risk, growth-versus-oversight ratios — are sector-specific additions Chapter 65 explicitly calls for, and Bimal realised he had been treating that guidance as an optional footnote rather than something he actually applied with real weight when scoring microfinance names. The calibration exercise made the cost of that neglect concrete: a footnote-level adjustment, never actually opened, could not have caught a tenfold difference in eventual NPL outcomes between two similarly-scored institutions.

KEY CONCEPT A dimension can fail calibration in two opposite directions — by carrying the right weight but being measured on the wrong time window for a seasonal sector (too loud, reacting to the calendar), or by carrying the right weight in principle while its sector-specific questions are never actually asked (too quiet where it matters most) — and both failures require a different fix.

The third pattern is the most dangerous, because it feels like careful, honest work while actually being its opposite: quietly reshaping the sector-specific questions and thresholds, case by case, after already knowing how each test case turned out, until the score would have gotten every single historical case right. This is overfitting wearing the costume of diligence — the same trap Chapter 79 named when it warned against tuning a trading rule until it perfectly matches the very data used to build it. If Bimal had simply invented an ever-more-specific set of MFI red flags and hydropower seasonal corrections, hand-tailored until his four test cases scored perfectly, he would have produced a process that explained the past flawlessly and would very likely fail the next situation Nepal's market produced, because it had been shaped to fit noise specific to these four cases rather than the genuine, durable guidance Chapter 65 already provides in general terms. The tell-tale sign of this trap, Bimal noted, is a calibration session that ends with the investor feeling triumphant rather than sober — real calibration, done honestly, should leave you a little uneasy, aware of how thin your evidence base still is.

CAUTION If a round of calibration ends with every single historical test case scoring perfectly, that is a warning sign of overfitting, not proof of a well-built score — genuine calibration usually leaves at least one case only partly explained.

Lesson 80.5 — Adjusting Responsibly: A Mandatory Checklist, a Written Log, and a Fresh Test

Having named the failure patterns, Bimal resisted an urge that surprised him: the temptation to "fix" the problem by simply moving points between Chapter 64's seven dimensions. He caught himself starting to sketch exactly that — trim Liquidity, add the difference to Financial Strength — before realising it would have been the wrong lesson entirely. Chapter 64's 100-point architecture had not failed either test case. The NIC Asia merger, the Chhimek comparison, and the hydropower quarter had all been mis-scored not because the seven dimensions carried the wrong number of points, but because Bimal had not been applying Chapter 65's sector-specific version of those dimensions with any real discipline. Moving points around would have quietly changed a book-wide standard — the same 20/15/10/15/15/15/10 architecture every other chapter in this Canon, including every case study, relies on — to paper over what was actually a gap in his own process. So he made no changes to Chapter 64's weights at all. Instead, he made two process changes, each tied to a specific piece of evidence, and each written down in a running log with the date, the reasoning, and the evidence that triggered it — the same kind of paper trail an NRB bank examiner would have insisted on before approving any change to a supervisory rating model.

The first change: for any hydropower company already past its Commercial Operation Date, Bimal committed to never scoring Financial Strength & Profitability off a single quarter's raw figures again. Every ROE and earnings-quality reading for a post-COD hydropower stock would now be computed either as a trailing-twelve-month figure or compared against the same quarter one year earlier, exactly as Chapter 65 instructs, before any point value was assigned. Reason logged: the dry-monsoon test case showed a single-quarter reading swinging the score by ten to fifteen points for reasons that had nothing to do with the company's underlying quality, purely because the seasonal comparison Chapter 65 already specifies had not actually been built into his process.

The second change: for any microfinance institution, Bimal committed to two mandatory questions before finalising Governance & Promoter Behaviour — has the institution's loan growth or NPL trend shown signs consistent with loan recycling, and is there evidence of borrower overlap with other lenders in the same operating district — and to explicitly weighting loan-loss provisioning coverage, not generic leverage, within Financial Strength & Profitability's capital-adequacy sub-component. Reason logged: the Chhimek-versus-weak-MFI comparison showed that these exact signals, already named in Chapter 65, were the single clearest predictor of which institution would go on to suffer severe NPL stress — and neither had actually been checked before 2021.

The third change, the smallest and most structural: Bimal built a one-page checklist, one row per sector, that he now pulls out and physically checks off before finalising any Canon Score — a forcing function to make sure Chapter 65's guidance gets applied every time, rather than trusted to memory. Reason logged: both failures traced back to the same root cause — sector-specific guidance that existed in the book but was never actually consulted at the moment of scoring.

Sector / situationWhat Bimal had been doingWhat Chapter 65 actually requiresProcess change logged
Hydropower, post-COD, Financial Strength & ProfitabilityScoring ROE and earnings quality off the most recent single reported quarterCompare the same quarter year-on-year, or use trailing-twelve-month figures, to strip out wet/dry seasonalityAdded a mandatory TTM/YoY check before finalising this dimension for any hydropower company
Microfinance, Governance & Promoter BehaviourGeneric checklist only — promoter pledging, related-party transactions, disclosure timelinessAlso check for loan recycling and multiple-borrowing/overlap risk specific to MFIsAdded two mandatory MFI-specific questions before finalising this dimension
Microfinance, Financial Strength & ProfitabilityGeneric leverage-style capital-adequacy sub-checkWeight loan-loss provisioning coverage more heavily than generic leverage, given unsecured group lending's structurally higher default riskReplaced the generic leverage sub-check with a provisioning-coverage-weighted version for any MFI

Chapter 64's 100-point architecture is untouched by any of this — no dimension gained or lost a single point. What changed was whether Chapter 65's already-written sector-specific instructions were actually being opened and applied at the moment of scoring, every time, rather than approximated from memory or skipped under time pressure.

PRACTICAL TOOL Keep a running, dated log of every process change made to how a scoring model is applied: what changed, which sector it affects, and which specific piece of evidence triggered it. A process change with no logged evidence behind it is indistinguishable, a year later, from a guess — and a scoring model's weights are not the only thing that can drift out of calibration; how faithfully its existing rules are actually applied can drift too.

The final and most important part of Bimal's method was what Chapter 79 called out-of-sample testing, applied here to a scoring process instead of a trading rule. Having changed his process based on the NIC Asia, Chhimek, weak-MFI, and hydropower cases, he did not declare victory. Instead he pulled a fifth situation he had deliberately set aside and had not looked at while making the changes — a small development bank that had gone through a rocky capital-raising period earlier in the decade — and scored it fresh using the new, checklist-enforced process. Only after checking that the new process still produced sensible, defensible judgments about a case that had played no role in shaping it did he consider the round of calibration provisionally complete. This is the same discipline as testing a trading rule on a different time period than the one used to build it: the evidence used to change the process and the evidence used to check the change must never be the same evidence, or the check proves nothing except that the process now agrees with itself.

Lesson 80.6 — What a Calibrated Score Can and Cannot Promise

After this exercise, Bimal's Canon Score was, in a real sense, better than it had been — not because it now guaranteed correct calls, and not because a single point had moved between dimensions, but because its application now reflected actual evidence about what has mattered for NEPSE outcomes rather than sector guidance that existed on paper but was not consistently opened in practice. That is a genuine and valuable improvement, and it is worth being precise about exactly what kind of improvement it is.

A calibrated score can promise that its application is no longer arbitrary — every sector-specific check it now relies on traces back to a specific, checkable piece of NEPSE history, and that history is written down in a checklist, not just remembered. It can promise a more honest starting point for the next stock an investor considers, one less likely to be silently misled by a dimension scored on the wrong time window, or blind to a sector-specific danger that Chapter 65 already named but that never actually got checked. It can promise that when the score turns out to be wrong about a future company, the investor will have a clearer basis for asking which dimension failed and why, rather than throwing out the whole system in frustration.

A calibrated score cannot promise that it will keep working forever without further attention. Nepal's capital market is still young and still thin — 271 listed companies is a modest universe by any regional standard, its sectors are lopsided and correlated in the ways Lesson 80.2 described, and the specific historical episodes available to calibrate against will keep being a small, imperfect sample for years to come. New kinds of companies will list, new regulatory regimes will reshape entire sectors overnight the way past NRB merger waves reshaped banking, and dimensions that look well-calibrated today may prove hollow the next time Nepal's market produces a genuinely new kind of stress. Calibration, done honestly, is not a task an investor finishes once. It is a practice an investor returns to, the same way Bimal, in his NRB years, never considered a bank's risk rating a permanently settled question — only a current best estimate, due for review the next time meaningful new evidence arrived.

Chapter recap

This chapter asked a narrower and more technical question than Chapter 79's general introduction to backtesting: not whether the discipline of testing an idea against history is worthwhile, but whether Chapter 64's seven real dimensions — Financial Strength & Profitability at 20 points, Governance & Promoter Behaviour, Valuation Reasonableness, Sector & Business Model Durability, and Growth Trajectory at 15 points each, and Liquidity & Tradability and Dividend & Capital Return Discipline at 10 points each — were actually being applied with Chapter 65's sector-specific rulers in practice, dimension by dimension, sector by sector, or merely being nodded at in principle. Calibrating a score this way is different from backtesting a trading rule: it opens up the recipe and tastes each ingredient separately, rather than only checking whether the finished dish came out well.

NEPSE's own structure makes this harder than it first appears. With roughly 271 listed companies clustered heavily into hydropower, microfinance, and banking, a calibration exercise that treats every company as an independent test case is fooling itself — most of those companies move together, reacting to the same monsoon cycle, the same NRB policy shift, or the same rural credit downturn, which means the real number of independent scenarios behind any NEPSE-wide test is a small handful, not hundreds. Every conclusion drawn from calibration has to be held with that humility built in.

Using real, verifiable NEPSE history — the 2013 merger that created NIC Asia Bank, the sharp divergence between strong and weak microfinance institutions during the 2021-2026 NPL stress episode, and a hydropower company's temporary, seasonally-driven score swing during a dry monsoon quarter — Bimal Sharma tested the Canon Score's pre-outcome verdicts against what actually happened. He found the score correct on the bank merger, right for an incomplete reason on the strongest microfinance institution, and outright wrong in failing to separate that strong institution from a weak peer that later suffered severe loan stress. That last finding was a genuine calibration failure, not a hypothetical one, and naming it honestly — rather than explaining it away — was the point of the exercise.

From these failures, the chapter drew out two opposite calibration problems worth watching for in any scoring system: a dimension scored on the wrong time window for a seasonal sector, so it sounds important but barely discriminates because it moves in lockstep across the entire sector at once (Financial Strength & Profitability, scored off a single raw quarter, in hydropower), and a dimension whose sector-specific questions exist in writing but are never actually asked at the moment of scoring (Governance & Promoter Behaviour's loan-recycling and overlap checks, and Financial Strength's provisioning-coverage weighting, in microfinance). It also named the trap that lurks behind both: quietly reshaping the scoring process after seeing outcomes until it explains history perfectly, which is overfitting dressed as diligence, not honest calibration. The responsible alternative demonstrated here left Chapter 64's 100-point architecture completely untouched — no dimension gained or lost a single point — and instead built a logged, evidence-triggered checklist that forces Chapter 65's sector-specific guidance to actually be opened and applied every time: a mandatory trailing-twelve-month or year-on-year comparison for post-COD hydropower companies, and two mandatory microfinance-specific questions for Governance plus a provisioning-weighted read of Financial Strength for lenders. That revised process was then tested against a fresh case that played no role in shaping it, echoing Chapter 79's insistence on keeping in-sample evidence and out-of-sample checks strictly separate.

A well-calibrated Canon Score, the chapter closed by arguing, becomes a more honest reflection of what has actually mattered for NEPSE outcomes so far — not a guarantee of future accuracy, and not a task that is ever finally finished, given how young and thin Nepal's capital market still is. That same unfinished quality applies just as much to the other half of any investing system: knowing when to sell. Chapter 81, "Testing and Refining Exit Rules," turns this exact discipline — real historical test cases, honest identification of failure patterns, small logged adjustments, and out-of-sample re-checking — away from the entry score covered here and onto the rules that govern when a Nepali investor should exit a position: stop-losses, profit-taking triggers, and the harder judgment calls about when a thesis has genuinely broken versus when it is merely being tested by short-term noise.

Primary data sources Figures, rates and rules referenced in this chapter can be verified against the primary sources: Nepal Rastra Bank (monetary policy, credit and BFI data), SEBON (regulation and issue approvals), NEPSE (prices, indices and turnover), CDSC (settlement and demat data) and Inland Revenue Department (tax rates and rulings). If a figure here disagrees with the primary source, trust the primary source and tell me.