✓KymiraCertified BI

THE DOCTRINE · PUBLISHED IN FULL

Eleven principles for numbers you can stake your name on.

This is the exact document that ships in every tier, the free one included. Not a summary, not a teaser: the product, demonstrating itself. Every skill you can buy exists to make an agent obey these eleven principles without being asked.

Eleven principles for building business intelligence you can trust. None of them is theory. Each one is the lesson left behind by a real wrong number that once reached a real page, and each one is now something the Kymira skills teach your agent to do without being asked.

A general coding agent does not show up believing these. That gap is the whole product.


1. An AI never authors a number

The model writes pipelines, layouts, and prose. Plain, inspectable code computes every figure that appears on a page. No exceptions, and no "it is only a summary stat."

"Page" means every shipped artifact. A figure appearing in the manifest, the handoff summary, a README, or a code comment the customer will read is either generated by the same code that computes the page, or gated by a check that scans prose files for numeric claims and fails naming the file and line. A manifest that contradicts its page is a failed build. The fastest guard is also the crudest: before shipping, grep the prose artifacts for digits that appear nowhere in the machine output.

This is what separates a certified report from a text-to-SQL demo, which routinely produces confident wrong answers. It is a rule about how the work is built, not a request to the model, so the model cannot be talked out of it.

Every computed number reconciles against the source file's own totals, not against your recomputation of them. If a file states no total to tie to, that is something to raise, not a check to quietly skip.

The honest limit, which you should state whenever you claim this principle: the model does not author the number, but it does author the plan. A wrong column or a wrong filter produces a wrong number that reconciles perfectly. The gates here cover arithmetic and identity. They do not cover meaning. That is why principle 10 exists, and why confirming a definition is not a formality.

2. Two independent reads beat one careful one

Find each value two different ways that share no logic: once by where it sits in the file, once by the label next to it. Both must agree, within a tolerance you write down.

A silent layout shift or a wrong-column read cannot fool two different methods the same way. When they disagree, you investigate. You never auto-resolve the disagreement, and you never widen the tolerance to make it go away.

3. Refuse to guess

When a label appears zero times, or more than once, that is an error. When a date will not parse, that is an error. When an expected column is missing, that is an error.

Guessing your way past ambiguity is exactly how wrong numbers get published. The correct response to ambiguity is to stop and say precisely what was ambiguous.

A date that parses two valid ways is as much an error as one that parses none. 4/9/2026 is April 9th or September 4th depending on a locale the file never states. Prove the format from rows the file itself disambiguates (any date field exceeding 12), apply that proof to the ambiguous rows, and disclose the proof on the page. If no row settles it, stop; publishing under a guessed locale is authoring a number.

4. Filenames lie

Vendor exports are named by the system that produced them, not by what they hold. A file called weekly_sales.csv may contain last week, another account, or both.

Verify identity and period from inside the file: pre-header rows, embedded titles, report metadata. Where a file carries no internal evidence of what it is, say so, and never compare a filename-derived value against itself, which is a test that always passes and proves nothing.

5. Fail closed and stale

A failed check blocks the update. The last known-good version stays live, marked so a reader knows a newer cut is under review, and a person is told.

Never publish a number you cannot stand behind. Stale and correct beats fresh and wrong, and the gap between them is where trust is lost for good.

Failing closed is an artifact, not an exit code. On any blocking failure, structural or arithmetic, first run or rerun, the build still writes a BLOCKED report page and machine summary carrying every check verdict earned before the stop, the named reason for the stop, and no certified figure, then exits non-zero. A bare abort that leaves only a log line, a stack trace, or an empty folder is itself a failed check.

Seen in our own adversarial audits: one ambiguous date in an appended row crashed a build with a raw ValueError and zero artifacts, on a file whose other 537 rows had already earned fourteen passing verdicts. The customer in that story gets a stack trace instead of the one page that says "blocked on row 538, here is everything the file still proves." The law exists so the second story is the only possible one.

And failing closed protects what already shipped: a build assembles every artifact in a staging directory and swaps it into the deliverable folder only after the final gate passes, so a failed run leaves the prior pack byte-for-byte untouched and visibly marked as having a newer cut under review.

6. Test that a check fires, not that it exists

The most dangerous check is one that can never trip. Every guard needs a test that feeds it known-bad input and confirms it blocks, not a test that feeds it good input and confirms nothing happens.

A check that cannot fail is a comment with a runtime cost. More than once, checks that looked green for weeks turned out to be structurally incapable of firing.

The fire test must call the deployed guard through the production code path and watch it block. A test that re-implements the guard's comparison inline can stay green forever while the real guard is broken; that is fire-test theater, and it fails this principle even though a test named after the guard exists and passes. A shipped check that structurally cannot fail is not a check, it is decoration, and rendering it beside real gates is a false claim.

Our own adversarial audits caught every shape of this in one round: a check reading "pass" if True else "fail" rendered green on four surfaces; a "two independent rails" verification that built both rails from the same field list and compared 537 values to themselves; five check rows whose text no input could ever change. Every one sat beside real gates, indistinguishable to the reader.

The enforcement is per row: every check row rendered on a page must have a fire test that turns that exact row red through the production pipeline, and a row without one may not render as PASS. The same law covers configuration: a knob that nothing reads is deleted, not shipped, because an inert setting is a lie in config form (principle 8's empty-config rule, applied to dead wiring).

7. Clean by relabelling, never by revaluing

Cleaning the data is where most of the real work lives, and where trustworthy reports usually die, because fixing the data and changing the numbers look identical from inside the code.

The line: you may change what a row is called. You may never change what a row says.

Fair game, because these concern identity: normalizing a name or SKU so two spellings resolve to one thing, unifying date or currency formats where the value is unchanged, removing proven-duplicate rows, classifying a row into a category, joining a row to a reference table.

Off limits, because these author values: filling in a missing number, smoothing or clipping an outlier, back-filling from a prior period, or deriving a value the source does not contain and presenting it as if it came from the source.

A missing number stays missing and is reported as missing. Blank is a fact about the source, and replacing it with a plausible figure destroys the one signal that something upstream broke.

This is testable, which is what makes it real rather than aspirational: after cleaning, the anchor must still tie. Relabelling cannot move a sum. If a total shifted while you were cleaning, something was revalued.

None of this forbids computing what a suspect value would mean. A flag without stakes is not a flag: every flagged figure states its share of each headline it feeds, and a labelled what-if value computed under each open ruling branch, each branch valued under its own hypothesis (an entry-error branch reprices the row; it does not delete it). A conditional figure computed by code and labelled as conditional is disclosure, not revaluing. A named doubt without its cost is a disclaimer, not a doubt.

When two or more rulings are open at once, value every combination of open branches, the full cross-product, each cell computed by code and labelled, so the reader's most likely question, the corrected figure under the adopted scope, is answered on the page and never left as subtraction for the reader.

8. Thresholds are configuration, not code

Every tolerance, bound, alias list, and content marker is a setting, not a magic number buried in logic. Code holds the mechanism, configuration holds the judgment. That line is what decides whether the second customer is an afternoon or a rebuild.

Empty configuration must fail loudly, never default to permissive. A check that silently passes because nobody configured it is worse than no check, because it reports a safety it never provided.

9. A platform's own numbers are claims, not facts

The moment a report ingests marketing data from Meta, Google, TikTok, or an analytics tool, it meets a kind of number that looks like a measurement and is actually an assertion by an interested party.

Meta reports the conversions Meta believes it caused, on Meta's window and model. Google reports Google's. Add them up and they routinely claim more revenue than the business actually made. These are not errors. Each platform is answering its own question in its own favor.

So record them as claims, name the claimant, and never blend them into a reconciled figure.

The same holds for any attribution your own report models. That output is your claim, labelled as yours, and it is never laundered into the reconciled column.

10. A metric without its definition is not a number

Across the metrics that consumer and operations teams actually use, competing definitions in live production are the norm, not the exception. Not most of them. Nearly all of them.

The worst offenders are also the most requested reports:

Why this is doctrine and not a documentation footnote: every one of these alternatives is internally consistent. Each reconciles perfectly against its own source and passes every gate above. Principle 1 stops a model from computing a figure. It does nothing about a figure meaning something other than what the reader assumes. This is the largest source of silent wrong numbers in the domain, and it is invisible to arithmetic.

So:

  1. Resolve the fork before computing. Ask in business language, with the consequence stated: "Percent of what, of what we shipped this period, or of what was on the shelf at the start?"
  2. Record the answer as a ruling: durable, attributed, reviewable, reversible.
  3. Show the definition beside the number, not in a tooltip nobody opens. Sell-through 62% (units sold / units received, 4-week).
  4. Never carry a definition across sources. Two retailers can mean different things by the same word, and a portfolio view that averages them silently is a number that describes nothing.

The Metric Library is the catalogue of these forks and the exact question each one needs. It is the most valuable reference after this document.

When no user can be asked, unattended, scheduled, or one-pass runs, neither guess nor refuse silently. Record a PROPOSED ruling: status proposed, owner unassigned, the fork, both readings with their numbers, and the provisional reading chosen. Publish behind a visible PROPOSED flag. Never attribute a decision to a person who did not make it; a ruling converts to active only when a named human ratifies it with their own timestamp.

11. Reconciliation proves arithmetic, not plausibility

A total that ties to the cent can still be 100x wrong. The fat-fingered price multiplies through every sum consistently, so every reconciliation passes while the headline carries a number no human would believe on sight.

Before certifying any figure, run a plausibility screen over every value column: the price, rate, and amount columns, and just as much the quantity columns (seats, units, counts), since a fat-fingered quantity certifies exactly as cleanly as a fat-fingered price. Compare each value against its group's typical value (modal or median per plan, SKU, or category), in both directions, high side and low side. Flag any row beyond the configured multiple (a threshold, per principle 8), state each flagged row's share of every headline it feeds, and say on the page that the screen ran and what it found.

When a quantity column's per-group typical value is degenerate against a legitimate discrete tier set, screen each value against the highest legitimate tier of its group with an exclusive boundary, publish the tier basis beside the threshold on the page, and record any unguided calibration as a PROPOSED ruling. The tier basis must never be derived from the candidate value itself: a ceiling that a slip can raise is circular, because the slip becomes its own ceiling and certifies clean.

Non-circular does not mean unavailable, and refusing to screen is not a neutral outcome. There are three legitimate bases, in order of preference: a configured tier list, the Metric Library or another stated outside prior, and, always available, a leave-one-out basis: compare each value against its group's distribution computed with that value excluded. A single slipped row cannot move a median it is not part of, so leave-one-out is non-circular by construction and needs no configuration. Only when a group is too small for a leave-one-out read (fewer than four other rows) may the screen report UNCALIBRATED for that group, and it reports which groups, not the whole column.

Be precise about what UNCALIBRATED costs, because an honest refusal to screen is still an unscreened column: the page states which columns and groups went unscreened, and no figure resting on an unscreened column may be described as certified. A screen that refuses everywhere and a screen that runs and finds nothing must never render the same.

A flag raised by any screen travels the full disclosure path or the build fails: the stakes, the share of each headline, the PROPOSED ruling, the what-if, and the flag glyph at every render of the affected figure. Detection that the render step then drops is worse than no detection: a check table that counts a flag the page denies is a manifest contradicting its page (principle 1), and it fails the build. In our audits the probe that proved this was an internally consistent 100x seats slip: the screen caught it, the check table counted it, and the certified page still showed the inflated headline at face value. Detection happened; the customer never saw it.

This principle exists because it was once the only one missing: in testing, every agent that caught a planted 100x slip caught it on its own initiative, and initiative is not a guarantee. Now it is written down.

It carries a second lesson, learned when the first fix failed. Closing the circular ceiling without naming a workable substitute produced screens that refused honestly and caught nothing, and a quantity slip certified as cleanly as before. Repairing the mechanism a rule names, while the harm the rule exists to prevent still lands, is not a fix. Test every rule against the harm, not against its own wording.


How they compose

Principles 1, 2, 7, 9, and 10 govern what a number is: derived not authored, cross-checked, never altered in transit, never confused with someone's claim about it, and never shown without the definition that gives it meaning.

Principles 3, 4, 8, and 11 keep the system honest about the limits of what it can prove. Principle 5 makes the failure mode safe and visible. Principle 6 makes the guarantees real rather than merely written down.

State the claim precisely, because it is narrow and it is the whole point. Not "the numbers are always right." Rather: no figure is authored by a model, errors fail loud and stale instead of silent and wrong, and every number traces back to the source file it came from. For a wrong figure to reach you, a bug would have to defeat the entire net above, and the net exists to catch bugs.


What a certified deliverable is

The principles govern the numbers; this governs the thing that carries them. The certified deliverable is a rendered, self-contained HTML page plus its machine artifacts: the data, the checks log, and the manifest. The definition sits beside every number, the check results are visible on the page, and every sibling artifact is linked from it. Markdown alone is a draft, not a deliverable; a buyer forwards a page, not a repository. In a standalone build, "intranet" means the deliverable folder, and every linking rule still applies.

Hygiene is part of the contract. Every emitted HTML artifact opens with <!doctype html> and <meta charset="utf-8"> in its first bytes, and the render check fails if any character on the page renders as a replacement glyph or a mojibake sequence. A deliverable that dies on first forward, a page that opens garbled, a PDF stamped with a build timestamp or a local file path, undoes every check that passed inside it.

The pack ships the mechanical half of this contract as code: gate/doctrine_gate.py (installed to ~/.kymira/gate/). Run it over the deliverable folder before certifying anything; a red gate is a failed build. It enforces what a script can prove: hygiene, leaked paths, prose figures against machine artifacts, binary-verified hashes, and fire-test evidence for every rendered PASS (its checks-log convention is documented in the script). The gate is itself fire-tested against nine violation classes. A green gate is necessary, never sufficient: it cannot judge meaning, only mechanics, so it ends where the principles above begin.

Three conventions make the contract auditable. The gate's verdict is captured inside the pack: run the gate with --stamp and it writes gate_verdict.txt into the deliverable (the one file no manifest lists or hashes); a pack with no captured verdict fails the plain gate run. "Every render of the affected figure" includes aggregates: a flagged row's marker travels to every rollup, region, month, and total that contains it. And check status is a three-word vocabulary: PASS (red-capable, fire-tested through the production path), INFO (informational, carries no green mark and claims no guarantee), and BLOCKED. A row that cannot fail may render as INFO or not at all, never as PASS; a self-referential check that resolves after print renders as INFO with its convention stated.

Get the doctrine + the free skill → See it obeyed, live Your first week with it