The Canon Data Pipeline
First published 26 Aug 2026 · Last verified 29 Aug 2026
Sunita Rai keeps a battered Lenovo laptop on her desk at the structural engineering firm where she works in Kathmandu, and for the first four years she owned shares, that laptop was also her entire investment record. Broker statements sat in her email inbox, unsorted. Annual reports lived wherever she'd last downloaded them, if she'd downloaded them at all. When her accountant asked for the cost basis on a rights allotment from three years earlier, she spent an entire Saturday searching her Meroshare login, her Gmail, and a shoebox of paper TMS slips before she found it. That Saturday is when she decided that if she was going to run her portfolio like an investor rather than a gambler, she needed a system for the data itself, not just for the decisions she made with it. The checklists in Chapter 101, the financial models in Chapter 102, the position-sizing and tracking sheet in Chapter 103, and the XIRR log in Chapter 104 are only as good as what feeds them. Every number in those tools comes from somewhere. This chapter is about knowing where, and about building the unglamorous plumbing that gets the right number into the right cell without eating your weekends.
Sunita's system, three years on, is nothing exotic. It is a folder structure on her laptop, mirrored to a cloud drive, a short weekly routine, and a habit of writing the source next to every number she records. It took her a weekend to set up and now costs her perhaps twenty minutes a week plus a couple of focused hours each quarter. That is the standard this chapter holds itself to. If what follows takes you longer than that to run in practice, you have overbuilt it.
Lesson 105.1 — Where the Data Actually Lives
Before you can organise data, you need to know where it originates. Nepali retail investors often rely on secondhand sources — a Sharesansar headline, a Facebook investment group screenshot, a friend's WhatsApp forward — because the primary sources feel harder to reach. They are not, once you know the map. There are six places that matter, and each produces a different kind of number.
The Nepal Stock Exchange's own website is the first stop, and it is more useful than most investors give it credit for. NEPSE publishes daily trading reports covering turnover, traded shares, transaction counts, and the day's price movements across every listed scrip, alongside floor sheets that show the actual trade-by-trade tape. If you want the closing price history for a stock, the daily turnover for the market as a whole, or the exact volume that changed hands on a given day, NEPSE's site is the primary record, not a news aggregator's summary of it. NEPSE also publishes corporate disclosures directly, since listed companies are required to file material information with the exchange, including everything from board meeting notices to dividend declarations to auditor's reports.
The Securities Board of Nepal, SEBON, is the regulator, and its site is where you go for a different class of document: prospectuses for public issues, merchant banker reports, SEBON directives and circulars that change the rules investors operate under, and enforcement actions or penalties against companies or brokers. If a company changes its capital structure, issues rights shares, or gets flagged for a compliance lapse, the notice usually shows up at SEBON before it becomes market gossip.
Meroshare, run through CDSC, is your own transaction ledger and the closest thing to a single source of truth for what you personally hold. It carries your current holdings, your historical transaction record including IPO and rights allotments, and your dematerialized share balance. If a broker's statement ever seems to disagree with your own memory of what you bought, Meroshare is usually the tiebreaker, because it reflects the depository record rather than any one intermediary's bookkeeping.
Your broker's TMS, the Trading Management System each brokerage runs, is where your buy and sell orders actually execute, and it produces the transaction statements and contract notes that carry brokerage fees, SEBON fees, and the exact settlement amounts. These numbers are what you need for computing your real cost basis and your realised gains, since they include the fee drag that a simple price-times-quantity calculation misses.
Nepal Rastra Bank's website is where the macro backdrop comes from: policy rate decisions, the monetary policy statement issued each year, directives to banks and financial institutions, remittance and BOP data, and inflation figures. None of this tells you what to do with a specific stock, but it shapes the environment every stock trades in, and the position-sizing judgment in Chapter 103 leans on it.
Finally, company annual and quarterly reports, usually PDFs posted on the company's own website or filed through NEPSE, are the primary financial statements: balance sheet, profit and loss, cash flow, and the notes that explain restatements, related-party transactions, and provisioning. This is the document of record for the financial modelling work in Chapter 102. A media summary of an annual report is not the annual report.
| Data Type | Primary Source | What You Get |
|---|---|---|
| Daily prices and turnover | NEPSE website | Closing prices, volume, floor sheet |
| Corporate disclosures | NEPSE, company website | AGM notices, dividend declarations, material events |
| Regulatory filings and directives | SEBON website | Prospectuses, circulars, enforcement notices |
| Your own holdings and transaction history | Meroshare, CDSC | Demat balance, allotment history |
| Trade execution and fees | Broker TMS statements | Contract notes, brokerage and SEBON fee breakdown |
| Macro data and policy | NRB website | Policy rate, monetary policy statement, inflation, remittance |
| Audited financials | Company annual and quarterly reports (PDF) | Balance sheet, P&L, cash flow, notes and restatements |
Sunita's early mistake, in her first year of investing, was treating Sharesansar and Merolagani as if they were primary sources. They are useful aggregators and she still reads them daily, but she now treats every figure she sees there as a pointer to go verify, not as the number she writes into her spreadsheet. The habit costs almost nothing once it is automatic, and it is the single biggest reason her records have held up over time.
Lesson 105.2 — The Folder and Naming Convention That Survives Years
The reason Sunita lost an entire Saturday hunting for a three-year-old cost-basis document was not that she failed to save it. She had saved it, in a folder called "Documents," under a filename her phone had generated automatically when she scanned it. A data pipeline that cannot be searched two years later is not a pipeline, it is a pile. The fix is not complicated, but it has to be consistent, because a folder system you abandon after three months is worse than none at all — you will have some documents in the system and some not, and you won't remember which.
Sunita's structure, adapted here as a template, is a top-level folder called something like "Investing" with a small number of subfolders underneath it, organised by function rather than by date, because you will usually be looking for a thing, not a moment.
Under "Investing," she keeps a "Disclosures" folder, subdivided by company ticker, so each listed company she holds or watches gets its own folder — NABIL, NLIC, HRL, and so on. Inside each company folder, every downloaded document, whether an AGM notice, a dividend declaration, or a NEPSE disclosure, gets saved with a filename that starts with the date in year-month-day order, followed by the company ticker and a short description. A dividend notice from NABIL filed in mid-2026 becomes something like 2026-07-15_NABIL_dividend-declaration.pdf. The year-first date format is deliberate: sorted alphabetically, the folder automatically sorts chronologically too, so a company's entire disclosure history reads top to bottom in the order it happened, with no extra effort.
A second folder, "Annual Reports," holds one subfolder per company, with each year's annual and quarterly report saved the same way: 2025-Q4_HRL_annual-report.pdf, 2026-Q1_HRL_quarterly-report.pdf. Because fiscal years in Nepal run mid-Ashadh to mid-Ashadh, she notes both the Nepali fiscal year and the corresponding English year range in a small text file at the top of each company folder, since this is the detail she has found herself forgetting fastest — whether a report labelled FY 2081/82 corresponds to the period she's modelling in Chapter 102.
A third folder, "Broker Statements," holds TMS transaction statements and contract notes, organised by broker if she has used more than one, then by year. A fourth, "Meroshare Exports," holds periodic CSV or PDF exports of her CDSC holdings and transaction history, dated the same way. A fifth, "Macro," holds NRB monetary policy statements and any directive she has needed to reference, organised by year rather than by topic, since she rarely needs more than the current and prior year's macro documents at once.
| Folder | Example Subfolder | Example Filename |
|---|---|---|
| Disclosures | NABIL | 2026-07-15_NABIL_dividend-declaration.pdf |
| Annual Reports | HRL | 2025-Q4_HRL_annual-report.pdf |
| Broker Statements | Broker name, then year | 2026_NIC-Asia-Securities_contract-notes.pdf |
| Meroshare Exports | By year | 2026-06-30_meroshare_holdings-export.csv |
| Macro | By year | 2026_NRB_monetary-policy-statement.pdf |
The point of this system is not tidiness for its own sake. It is that the moment a discrepancy comes up — and it will, as the worked example in Lesson 105.4 shows — you need to go from "I remember there was a dividend notice about this" to the actual document in under a minute. A folder structure you have to think hard about to navigate fails exactly when you need it most, under time pressure, trying to settle a question before you decide whether to buy or sell.
One more discipline worth adopting alongside the folders: never rename a source document's content, only its filename, and never edit a PDF. If a company later restates a number, save the restated version as a new dated file rather than overwriting the old one. The old, wrong number is itself useful information — it tells you the company has a habit of restating, which belongs in your notes on that company's disclosure quality.
Lesson 105.3 — The Update Cadence: Daily, Weekly, Quarterly, Annual
The mistake that turns a data pipeline into a second job is trying to keep everything current all the time. NEPSE's floor sheet updates every trading day; you do not need to download it every trading day. The right cadence matches how often each tool in Chapters 101 through 104 actually consumes the data, not how often the data changes.
Daily, the only thing worth even glancing at is the closing price of stocks you hold or are actively watching for a buy, which takes under five minutes if you resist the urge to also read every headline and forum post attached to it. Sunita checks this from her phone on the bus home, using NEPSE's site or an aggregator, and does nothing else with it unless a price move is large enough to warrant checking whether a disclosure explains it.
Weekly, she spends about fifteen minutes updating the tracking sheet from Chapter 103: entering the week's closing prices, noting any dividends or corporate actions announced that week by checking the disclosure feeds for the handful of companies she holds, and glancing at NEPSE's weekly turnover summary to keep a feel for overall market activity. This is also when she scans for any SEBON circulars that might affect her holdings or her broker.
Quarterly, the workload steps up but stays bounded. When a company she holds files its quarterly report, she downloads it, saves it under the naming convention from Lesson 105.2, and updates the financial model inputs from Chapter 102 — revenue, net profit, EPS, and whatever ratios that model tracks. She also pulls a fresh Meroshare export of her holdings and transaction history as a reconciliation check against her own tracking sheet, and she reviews NRB's latest data release if one has come out in that window, since monetary policy shifts often land quarterly. This is the point where she also does her cross-checking pass, described in the next lesson, comparing what the company reported against what any news coverage said about it.
Annually, she does the heavier lift: downloading every annual report for every company she holds, re-verifying the XIRR calculation in Chapter 104 against a full-year Meroshare export, reviewing NRB's annual monetary policy statement in full rather than skimming it, and doing a full audit of her folder structure to catch anything that slipped through during the year, misfiled or missed. This is also when she prunes: companies she has fully exited get their folders archived rather than deleted, since she has more than once needed an old record for tax purposes.
| Cadence | What to Pull | Which Tool It Feeds |
|---|---|---|
| Daily | Closing prices for held or watched stocks | Chapter 103 tracking sheet, mental checkpoint only |
| Weekly | Price updates, new disclosures, SEBON circulars | Chapter 103 tracking sheet |
| Quarterly | Quarterly reports, Meroshare export, NRB data release | Chapter 102 financial model, Chapter 103 reconciliation |
| Annually | Annual reports, full Meroshare export, NRB annual statement, folder audit | Chapter 102 model refresh, Chapter 104 XIRR verification |
The discipline here is restraint as much as diligence. NEPSE, SEBON, and company sites will all happily give you more data than any four tools need. The cadence above is deliberately matched to the update frequency each of the earlier chapters' tools actually requires — there is no value in refreshing a quarterly EPS figure daily, and no value in letting a weekly price update slide for a month.
Lesson 105.4 — Cross-Checking and Data Quality
Numbers from Nepali capital market sources are not always clean on first pass, and treating any single source as automatically authoritative is a mistake, including this book's implicit ranking of primary over secondary sources. Companies restate prior-period figures, sometimes with little fanfare. PDFs get typos, especially in tables where a single misplaced decimal or an extra zero can survive multiple rounds of review. A news article summarising a quarterly result may be reporting a preliminary or unaudited figure and never issue a correction once the audited version differs. And straightforward transcription error, someone fat-fingering a number into a spreadsheet or an article, happens more often than anyone likes to admit.
That worked example is the entire argument for this chapter in miniature. The pipeline did not prevent the discrepancy from existing, discrepancies between unaudited and audited figures are routine in Nepali reporting, since companies often release preliminary results ahead of the full audited annual report, and the two can legitimately differ. What the pipeline did was make the discrepancy cheap to resolve, because both source documents were dated, filed under the same company ticker, and sitting one folder-click away instead of buried in an inbox or a phone's downloads folder.
The general habit worth building from this is threefold. First, always note whether a figure you are recording is unaudited, provisional, or audited final, since Nepali companies frequently release results in stages and only the audited annual report should be treated as settled. Second, when a figure from a news source and a figure from the company's own filing disagree, trust the filing, but keep both, dated, so you can explain the gap later rather than simply overwriting one with the other and losing the history. Third, treat any number that looks unusually clean, a suspiciously round figure, an outlier growth rate, a ratio that jumps sharply from the prior period, as a prompt to check the source PDF directly rather than the aggregator's rendering of it, since transcription errors compound when an aggregator scrapes a number and a dozen forum posts copy the aggregator.
There is a second, subtler quality problem worth naming: cross-source disagreement that has nothing to do with restatement, simply human error somewhere in the chain. A dividend percentage reported as 15 percent in one aggregator and 1.5 percent in another is not a restatement, it is almost certainly a typo somewhere, and the only way to resolve it fast is to go to the company's own disclosure or the NEPSE filing directly, which is exactly the habit Lesson 105.1 pushes toward. Investors who skip building direct-source habits end up needing to resolve this kind of disagreement from scratch every time it happens, at whatever inconvenient moment it surfaces, usually right when they are deciding whether to act on the number.
Lesson 105.5 — Backing Up What You've Built
A folder full of years of carefully dated, carefully named documents is worth exactly nothing the day a laptop's hard drive fails, and hard drives do fail, along with laptops getting stolen, dropped, or simply lost to a factory reset that nobody backed up first. This is not a hypothetical risk raised for the sake of thoroughness. It is one of the most common ways Nepali retail investors lose years of records, and it is entirely preventable with a habit that costs almost no ongoing effort once it is set up.
The minimum viable backup is a cloud sync of the entire "Investing" folder structure described in Lesson 105.2, using any mainstream cloud storage service, Google Drive, Dropbox, or a similar provider, configured to sync automatically rather than requiring a manual upload step. Automatic sync matters more than which provider you choose, because a backup that depends on remembering to do it manually will lapse exactly when life gets busy, which is also when you are least likely to notice a hard drive starting to fail. Sunita uses a folder that syncs continuously in the background, so a new disclosure she saves on a Tuesday evening is already backed up before she closes her laptop.
Beyond the folder of documents, the spreadsheet tools themselves, the tracking sheet from Chapter 103 and the XIRR log from Chapter 104, deserve their own explicit backup discipline, since these are working files you edit constantly rather than static documents you file once. If you keep these in a cloud-native spreadsheet tool, version history is usually built in and you get some protection automatically. If you keep them as local Excel files, the same automatic folder sync covers them, but it is worth occasionally saving a dated snapshot copy, a year-end copy renamed with the date, separately from the live working file, so that a spreadsheet formula error that silently corrupts months of entries does not also corrupt your only backup of the correct version.
A second layer worth considering, though not strictly necessary for most individual investors, is an occasional export to a second location entirely independent of your primary cloud provider, a periodic download to an external drive kept at a different physical location, or a second cloud account. This is insurance against the low-probability but non-zero case of an account lockout or provider-side data loss, and it costs perhaps an hour a year to maintain.
None of this is complicated, and that is deliberate. The investors who lose years of data are rarely undone by an exotic failure. They are undone by an ordinary one, a cracked screen, a stolen bag, a child who resets a tablet, on top of a backup habit that existed in intention but not in practice. Building the habit into the same rhythm as the quarterly and annual reviews already described means it happens automatically rather than needing to be remembered as a separate task.
Lesson 105.6 — Closing the Loop With the Rest of the Toolkit
The reason this chapter exists inside Part XVIII rather than standing alone is that a data pipeline is not useful in isolation. Its entire value is in how cleanly it hands off to the tools built in the four preceding chapters, and it is worth walking through that handoff explicitly so the whole Part reads as one system rather than four separate exercises plus a filing chore.
The pre-buy checklist from Chapter 101 draws directly on the Disclosures and Annual Reports folders: before initiating a new position, Sunita's checklist step of reviewing recent corporate actions and disclosure history is really just a prompt to open that company's folder and read the last year or two of filings in order, which the chronological filename sorting from Lesson 105.2 makes trivial. A company whose folder shows a pattern of late filings, frequent unaudited-to-audited restatements, or unexplained gaps is telling you something about its governance quality before you have looked at a single financial ratio.
The financial model in Chapter 102 consumes the audited annual and quarterly reports directly, and the discipline from Lesson 105.4, labelling every figure as unaudited or audited and preferring the audited version once available, is what keeps that model from being quietly built on a number that later gets restated out from under it. The quarterly update cadence from Lesson 105.3 is timed specifically so the model gets refreshed exactly when new audited or unaudited data becomes available, no faster and no slower.
The position-sizing and tracking sheet in Chapter 103 draws its price data from the daily and weekly cadence, and its transaction records from the Broker Statements and Meroshare Exports folders, reconciled quarterly. This reconciliation step, checking the tracking sheet's running record against a fresh Meroshare export, is where errors get caught before they compound. A transaction entered with the wrong quantity or price in week three will silently distort every subsequent calculation until someone checks it against the depository record, and quarterly is a reasonable frequency to catch that before it does real damage.
The XIRR log in Chapter 104 depends entirely on having a clean, complete, correctly dated transaction history, which is precisely what the Meroshare and broker statement folders are built to preserve. An XIRR calculation is only as trustworthy as its cash flow dates and amounts, and a data pipeline that has been quietly losing or misdating transactions for two years will produce an XIRR figure that looks precise and is actually wrong, with no obvious warning sign, since the formula will happily compute a confident-looking answer from bad inputs.
Seen this way, the six lessons in this chapter are really one lesson applied six times: know where a number actually comes from, keep it somewhere you can find it years later, refresh it only as often as it needs refreshing, verify it before trusting it, protect it from the ordinary disasters that erase unbacked data, and route it to the tool that needs it. None of the four preceding chapters' tools require anything more elaborate than that to function well. A checklist, a model, a tracking sheet, and an XIRR log are all, at bottom, simple structures that do not reward being fed by an overengineered data operation any more than a well-run kitchen rewards a walk-in freezer for a single-person household.
Sunita's system today looks almost exactly as described in this chapter: a folder tree that took a weekend to set up, a weekly fifteen-minute update, a quarterly afternoon of reconciliation, an annual review, and a cloud backup she tests once a year. It is not more elaborate than that, and it does not need to be. The temptation, especially for an investor who enjoys the mechanics of organising data, is to keep adding layers, more folders, more automation, more sources tracked "just in case." Resist that temptation the same way you would resist overcomplicating a position-sizing rule or a checklist. The pipeline exists to serve the four tools around it, not to become a fifth hobby competing with the actual work of picking and holding good investments.
Chapter recap
This chapter mapped the six places Nepali market data actually originates: NEPSE's own site for prices, turnover, and disclosures; SEBON for regulatory filings and directives; Meroshare and CDSC for your personal holdings and transaction history; broker TMS statements for execution detail and fees; NRB's site for macro data and policy; and company annual and quarterly reports for the audited financial statements everything else ultimately rests on. It laid out a folder-and-naming convention, organised by company ticker with year-first dated filenames, designed to remain searchable years after the fact rather than only in the week a document was saved. It set an update cadence, daily price glances, weekly tracking-sheet updates, quarterly report downloads and reconciliation, annual full reviews, matched to how often the tools in Chapters 101 through 104 actually need fresh input rather than to how often the underlying data changes. It walked through a real cross-checking failure mode, an unaudited EPS figure later restated in the audited annual report, and showed how a dated, well-filed paper trail turned what could have been hours of confusion into a ten-minute resolution. It laid out a backup discipline built around automatic cloud sync, annual restore testing, and dated snapshots of the working spreadsheets, aimed squarely at the ordinary, common way years of manually collected data get lost. And it closed by tracing how this pipeline feeds directly into the pre-buy checklist, the financial model, the tracking sheet, and the XIRR log built in the four preceding chapters, arguing that the pipeline's only job is to stay simple enough to serve those tools without becoming a project in its own right.
Chapter 106, The NEPSE Data Handbook — Historical Almanac, turns from process to reference. It compiles a compact historical almanac of NEPSE data, the kind of index-level and market-wide figures worth having on hand rather than re-deriving from scratch each time a question about NEPSE's history comes up, so that the pipeline built in this chapter has a settled body of historical reference material to sit alongside the investor's own ongoing records.