Share of 8,082 observed outcomes captured at the top 1, 5 and 10 percent investigation capacities

Methods study / Risk modelling / Norway

Point-in-time prediction of bankruptcy and forced dissolution

Development, adversarial audit and out-of-time validation of a 24-month risk model for Norwegian limited companies. The study shows strong ranking and substantial risk concentration, but also why calibration, date provenance and outcome maturation must be kept distinct from discrimination.

Methods study 8,736 words Norsk utgave
Study design 24 months from 31 December

The question is not whether a model can explain yesterday's bankruptcies. The question is whether it can rank tomorrow's company failures without having seen the future.

ROC-AUC0.882950Temporal FY2024 evaluation
PR-AUC0.411187Against prevalence 0.021513
Hits in top 1%2,352Of 8,082 observed failures
Paired bootstrap1,000Company-level draws, identical for both models
Established The model ranks later company failures strongly in the observed FY2024 cohort.

ROC-AUC 0.883, AP 0.411 and statistically supported gains over the frozen reference on the integrated ranking metrics.

Not established The score is not yet a final, calibrated probability of failure.

The 24-month outcome window is incomplete and calibration-in-the-large failed. The appropriate use is ranking and controlled shadow testing.

Abstract

Background: Credit-risk models can score well for the wrong reasons. Common failure modes include outcome-conditioned cohort selection, accounts that were not yet public at prediction time, overlapping outcome windows, contemporaneous registry snapshots, and fields that indirectly reveal whether a company was later deleted. We therefore built a modelling process in which temporal order and source provenance are treated as part of the statistical model itself, not as downstream data cleaning.

Methods: A proprietary, curated longitudinal reconstruction of publicly available Norwegian registry, annual-accounts and gazette data was anchored to a fixed prediction origin: 31 December of the accounting year. The outcome was bankruptcy opening or forced dissolution within the following 24 months. The model is an ensemble of ten gradient-boosted decision trees over 214 data-driven features plus one deterministic canary. Development and internal validation were separated at the company level; FY2024 served as a later out-of-time evaluation. The improvement over a frozen 198-feature reference was examined with 1,000 paired company-bootstrap draws.

Results: On 375,688 FY2024 companies with 8,082 observed failures, the ensemble reached ROC-AUC 0.8830, average precision (AP — in effect the area under the precision–recall curve, often labelled PR-AUC) 0.4112, standardised partial ROC-AUC at a false-positive rate of up to 10 percent of 0.7680, and a Brier score of 0.0161. The top 1, 5 and 10 percent of the ranking captured 29.10, 55.85 and 67.03 percent of the failures respectively. Against the reference, the difference in ROC-AUC was +0.0037 (95 percent confidence interval +0.0029 to +0.0047) and in AP +0.0132 (+0.0099 to +0.0163). Seed 42 replayed bit-exactly.

Interpretation: The evidence supports strong temporal ranking and a statistically robust incremental improvement. It does not yet support use as a finished absolute 24-month probability: the observed/expected ratio was 1.6216 — a lower bound until the window matures — and the pre-declared calibration-in-the-large gate failed. Moreover, the FY2024 outcome was observed up to 24 August 2026, while the full 24-month window closes on 31 December 2026. The study thus documents a strong development candidate for risk ranking and controlled shadow testing, not an unconditional basis for autonomous credit decisions.

Keywords: bankruptcy risk, forced dissolution, Norwegian limited companies, point-in-time, registry data, PR-AUC, calibration, company bootstrap, temporal validation.
Sammendrag (norsk)Les det norske sammendraget

Bakgrunn: Modeller for kredittrisiko kan oppnå høye skårer av feil grunner: utfallsbetinget kohortseleksjon, regnskap som ikke var offentlige på prediksjonstidspunktet, overlappende utfallsvinduer, samtidige registerbilder og felt som indirekte røper senere sletting. Vi bygde en modellprosess der tidsrekkefølge og kildeproveniens er en del av selve den statistiske modellen.

Metode: En kuratert longitudinell rekonstruksjon av offentlige norske register-, regnskaps- og kunngjøringsdata ble forankret i et fast prediksjonstidspunkt: 31. desember i regnskapsåret. Utfallet var konkursåpning eller tvangsoppløsning de neste 24 månedene. Modellen er et 10-seeds-ensemble av gradientforsterkede beslutningstrær med 214 datadrevne variabler pluss én deterministisk kanari. Utvikling og intern validering ble adskilt på selskapsnivå; FY2024 var en senere ut-av-tid-evaluering. Forbedringen mot en frosset 198-variabelreferanse ble undersøkt med 1 000 parrede selskapsbootstrap-trekk.

Resultater: På 375 688 FY2024-selskaper med 8 082 observerte utfall oppnådde ensemblet ROC-AUC 0,8830, PR-AUC 0,4112, standardisert partiell ROC-AUC (FPR ≤ 10 %) 0,7680 og Brier-skår 0,0161. Topp 1, 5 og 10 prosent av rangeringen fanget 29,10, 55,85 og 67,03 prosent av utfallene. Mot referansen var ΔROC-AUC +0,0037 (95 % KI +0,0029 til +0,0047) og ΔAP +0,0132 (+0,0099 til +0,0163). Seed 42 lot seg gjenskape bit-for-bit.

Fortolkning: Evidensen støtter sterk temporal rangering og en robust inkrementell forbedring, men ennå ikke bruk som ferdig absolutt 24-måneders sannsynlighet: O/E-forholdet var 1,62 og den forhåndsdefinerte kalibrering-i-det-store-porten feilet. FY2024-vinduet er høyresensurert til 31. desember 2026. Studien dokumenterer en sterk utviklingskandidat for risikorangering og kontrollert skyggetesting — ikke et ubetinget grunnlag for autonome kredittbeslutninger.

Read the complete Norwegian edition

1. Problem and scientific contribution

Bankruptcy is a rare, time-dependent and administratively recorded outcome. That makes the task harder than ordinary binary classification. A model can rank well because it genuinely detects economic deterioration — but also because it was shown a registry trace that only came into existence after the bankruptcy, because the hardest companies never entered the cohort, or because the same terminal event sits in both the training and the evaluation window. A high ROC-AUC is therefore not, by itself, proof that a model measures future risk.

Our study asks three questions. First: how well can a model rank bankruptcy or forced dissolution within 24 months using only information that was, in principle, available at a common historical cut-off? Second: does the result survive targeted attempts to falsify the signal — time gates, source prohibitions, duplicate controls, block shuffling and an exact replay? Third: does the ranking deliver practical concentration of failures when investigative capacity is limited to the top-ranked 1, 5 or 10 percent?

The contribution is first and foremost methodological. We treat data foundation, cohort, label and model as one connected measurement system. That matters because predictive performance is not merely a property of the algorithm; it is a property of the entire chain from public event to time-stamped feature to evaluated prediction. This reasoning follows the basic transparency of the TRIPOD framework: a prediction study must describe participants, predictors, outcome, validation and analysis precisely enough that the result's applicability and risk of bias can be assessed.

The study is not a comparison with banks or commercial credit-bureau products. We do not have access to their identical portfolios, outcome definitions, time points or decision costs, and such a claim would therefore not be verifiable. The results apply to the defined Norwegian AS/ASA cohort, the defined 24-month outcome and the frozen evaluation described here. The cohort comprises all Norwegian private and public limited companies (AS/ASA) that were at risk at the origin, not only small enterprises; the page's URL reflects an earlier working title.

2. Data foundation and historical outcome reconstruction

The data foundation is Byggsikt's own curated and quality-assured reconstruction of publicly available Norwegian registry, annual-accounts and gazette data. The sources comprise entity information, filed annual accounts and dated announcements from the Brønnøysund Register Centre. The collection infrastructure, the deterministic document interpretation and certain search procedures are proprietary. The scientific description nevertheless reports source category, temporal scope, logical de-duplication, transformation, row counts and cryptographic fingerprints for the model inputs.

The Brønnøysund Register Centre describes the Register of Company Accounts as a central source of insight into the financial condition of Norwegian business. Annual accounts are public, but they represent a reporting process: the accounting year closes, the accounts are adopted, filed and approved later. The ordinary deadline for avoiding late-filing penalties is 31 July of the following year. Accounts for FY can therefore not simply be treated as known on 31 December of FY. Concretely, only accounting years up to and including FY−1 enter as financial predictor information at the origin of 31 December of FY. Where an approval or filing timestamp is observed in the gazettes, it is additionally required to lie on or before the origin; for historical values without such a bound timestamp, temporal correctness rests on the FY−1 shift alone. The residual risk from missing document-level provenance is addressed in section 10.

The outcome was reconstructed from dated public announcements of bankruptcy and forced dissolution. The Bankruptcy Register's announcements contain, among other fields, the organisation number, opening date, case number and judicial venue, and are normally published on the day of the opening. We use event date and organisation number to establish terminal chronology. Voluntary liquidation, merger and demerger are treated as competing exits — neither as survival nor as bankruptcy. A company that already had a terminal outcome before the prediction date does not enter as an at-risk company at that date.

Historical reconstruction was necessary because contemporaneous business registers do not, on their own, represent the past's at-risk population. Companies deleted after bankruptcy may lack rich fields in a later registry extract. If the cohort requires such a field for inclusion, precisely many of the fastest and hardest bankruptcies are removed. We therefore made cohort entry independent of accounts availability. Missing observation is handled inside the feature process itself; it is not used as grounds for removing the company from the population.

Another problem is duplicated physical announcement rows. Count and recency features were therefore built after canonical de-duplication on logical event keys. A duplicate must neither move the date of the most recent event nor increase the event count. For families where the same day can contain conflicting states, a conservative unknown state is applied rather than choosing a favourable ordering that cannot be documented.

Source categoryTemporal roleUse in the studyPrincipal control
Entity registerHistorical identity and status informationCohort keys and stable company attributesNo contemporaneous "alive/deleted" fingerprints as model features
Annual accountsFY−1 or older at originLevels, quality and robust financial transformationsLagging, pre-origin evidence and source-adjusted shuffling
Company and accounts announcementsDate ≤ 31 Dec FYRecency, capital, auditor, roles and changesPhysical date gate, canonical de-duplication and future-append test
Bankruptcy announcementsDate > origin and ≤ origin + 24 monthsOutcome reconstruction onlyForbidden in predictor history

3. Estimand, cohort and time axis

The estimand is the probability that the company's first observed bankruptcy opening or forced dissolution occurs within the 24 months following 31 December of a given accounting year. The study primarily validates the relative ranking of that probability; the absolute level is assessed separately against pre-declared calibration gates in section 8, and did not pass. The prediction date is called the origin. Every feature must be traceable to information dated on or before the origin. A terminal event after the origin is an outcome; the same event can never simultaneously be history in the predictors.

Y(i,t) = 1 if Tfailure,i ∈ (31 Dec t, 31 Dec (t+2)] and no competing exit occurs first; otherwise Y(i,t) = 0 once the full observation window is known.

For FY2024 the window is not yet fully observed: companies with no recorded event as of 24 August 2026 are provisionally coded Y = 0 in this report. This right-censoring — an outcome that is unknown, not known to be negative — is treated explicitly in sections 8 and 10.

The cohort aims to represent limited companies that were at risk at the origin. It is therefore not conditioned on whether an annual account exists in a later extract. This point matters: a requirement of an accounts row can look neutral, but companies that fail early often never file their next accounts. "Has accounts" then becomes selection on future survival. In this study the company is retained, while unavailable values are handled by a fit-only process without any model-facing coverage or source flags.

Competing terminal exits are handled chronologically. If voluntary liquidation, merger or demerger occurs before bankruptcy or forced dissolution, the outcome is coded as not having occurred in the window (Y = 0): the competing exit precludes failure, and the company remains in the population. This is a competing-risk treatment of the estimand — the target is the cumulative incidence of failure as the first terminal event — not censoring; the word censoring is reserved for windows that are not yet fully observed, as with FY2024 above. If failure occurs first, the outcome is positive. The principle prevents a company from being described both as excluded and as later positive on the strength of an arbitrary priority between columns. The formally improved chronological label code produces a very small difference against the label version that generated the frozen result; this is reported as a reproducibility limitation later in the article.

Figure 1 / Temporal design

One clock for predictors, one later clock for outcomes

Point-in-time
A point-in-time design means more than removing obvious outcome fields. The entire data generation must respect the same origin. The final part of the FY2024 window is flagged as an explicit limitation, not treated as known negative outcomes.

Development, validation and evaluation roles

The frozen run used 300,863 companies for model fitting and 33,236 companies for internal validation; 765 validation companies had a positive outcome. The split was deterministic on organisation number, so that no company appeared on both sides of the fit/validation boundary. This prevents repeated company-years or stable company characteristics from leaking between the two roles. The FY2024 evaluation consisted of 375,688 companies with 8,082 recorded failures.

With a 24-month horizon, the outcome window of the most recent development vintage partially overlaps the FY2024 evaluation window: the development material extends through FY2023, whose window shares calendar year 2025 with FY2024's, so the same terminal event can serve both as a training label and as an evaluation outcome for the same company. The frozen run was produced with the earlier label vintage, which did not purge such overlapping rows; a purge was introduced only in the later, chronologically corrected label code discussed in section 10. The shared window must therefore be counted among the sources of optimism in this evaluation, alongside the limitations listed there. The row counts for the overlap belong in a future version of the evidence file before they are cited.

The FY2024 evaluation is also out-of-time, not out-of-company: a company that appears in the development material with earlier accounting years also enters the FY2024 cohort with its FY2024 observation. The company-disjoint split applies to the fit/validation boundary within the development material. The generalisation being tested is thus to a later time period for a largely overlapping population of companies, not to unseen companies.

FY2024 is temporally later than the development material and is therefore an out-of-time evaluation. It is nevertheless not an untouched external confirmation cohort. Several development choices were, over time, assessed against FY2024. The confidence intervals in this article therefore describe the uncertainty in the frozen, paired comparison; they do not include the full uncertainty of the earlier adaptive model search. A genuinely blind confirmation requires a new time cohort, or an independent portfolio, that is opened only after architecture, label and analysis plan are locked.

RoleCompaniesFailuresUse
Fit300,863Not published as a headline estimateTree fitting and fit-only imputation
Validation33,236765Early stopping and calibrator, fully company-disjoint
FY2024 evaluation375,6888,082Later temporal performance and paired comparison

4. Feature architecture: mechanisms rather than field hunting

The frozen candidate has 215 inputs: 214 data-driven features and one deterministic pseudorandom canary. The canary is an audit instrument, not economic information; it is a fixed, outcome-blind function of stable identifiers, so its value is bit-reproducible and by construction independent of the outcome. The data-driven features are organised into families with an explainable mechanism: payment capacity and capital buffer, accounts quality, change intensity and recency, company age, auditor history, industry, and historical connections between persons and other companies.

The accounts block is the largest. Level features preserve information that can vanish in ratios alone, while quality features describe consistency and structure in the available accounts observations. A separate control showed that recomputing 14 traditional financial ratios barely improved the trees; flexible trees could largely form the same splits from the levels. That does not mean accounts are unimportant. On the contrary, a source-adjusted shuffle of the historical accounts block showed that company-specific values carried substantial predictive information. The negative finding concerns redundant transformation, not the accounts source itself.

The recency features measure how many days have passed since particular types of dated announcements. Earlier experiments with a sequence model found no stable added value over a simple count-and-recency representation. The useful finding was that "days since" was consistently more informative than mere counts. The frozen architecture therefore uses a transparent, aggregated event history rather than retaining a volatile sequence model that never beat its own baseline.

Person and auditor networks can be strong, but they are also dangerous. A contemporaneous role snapshot taken years later reveals who was still active — and thereby, indirectly, who survived. We use only dated role events before the origin, identify persons within each historical point in time, and apply leave-one-out with respect to the company being scored. History based on bankruptcy openings is forbidden as a predictor, both because the outcome itself can leak and because public retention of bankruptcy history varies over the calendar. Network signals are instead built from dated, non-terminal warning events with better temporal coverage.

Industry is represented from point-in-time purpose text and official classification, without feeding the model's classifier confidence in as a separate feature. That last point is an important design choice: the confidence of a text classifier can measure how similar the text is to the population the classifier was trained on, and thus become a hidden survival signal. The candidate uses industry information, but not a margin that claims to know when the industry guess is "certain".

FamilyCountEconomic or administrative mechanism
Control and base features53*Cohort-stable base signals, industry representation and one deterministic canary
Event recency31Time since dated changes, warnings and administrative events
Historical person network9Pre-origin connections to other companies and warning history, leave-one-out
Accounts levels75Liquidity, capital structure, profitability, size and result
Accounts quality30Internal structure and observable quality of historical accounts values
Capital history9Level, change, minimum capital and dated capital events
Age and auditor core5Observed historical age, auditor change and explicit auditor clearance
Auditor network, provisional3Dated engagements and warning recency of other known clients

* The 53 control and base features include one deterministic canary. The candidate thus has 214 data-driven features, not 215 independent economic signals.

Missing accounts: an observation problem, not automatically an economic property

Missing accounts values can arise because the company did not file, because the public history is limited, because the document was retrieved through a different historical channel, or because a field is not applicable. If the model is given an explicit "source", "retrieved" or "recovered" flag, it can learn the data-collection process instead of credit risk. Such model-facing provenance and coverage flags have therefore been removed.

Numerical gaps are filled deterministically from donors selected in the fit population only, primarily within industry group, and without reading the label. Observed cells are preserved. This prevents evaluation data from influencing the imputation and makes the missingness treatment outcome-neutral. The method is nevertheless not equivalent to multiple imputation and does not express full uncertainty about each missing value. Robustness is therefore also assessed by shuffling and by comparison against an independent accounts panel on overlapping companies.

5. Model and training design

The model family is gradient-boosted decision trees for a binary outcome. Trees are well suited when financial levels, event recency and network signals carry thresholds and interactions that are not necessarily linear. The architecture was frozen before the final ten-seed run. Each member used 63 leaves, learning rate 0.05, 80 percent random feature sampling per iteration, 80 percent row sampling, at least 100 observations per leaf and a raw class weight of 1.0. Early stopping used a patience of 100 iterations without improvement in validation-set ROC-AUC as its stopping criterion; the best iteration for seed 42 was 320 (section 6). The hyperparameters were selected in the preceding development process described in section 3 — a process that was not blind to FY2024 — and then frozen before the ten-seed run. The uncertainty attached to that selection is not represented in the reported confidence intervals.

Seeds 42–51 varied the row and feature subsampling, while the data split and all other contracts were held constant. The ensemble prediction is the mean of the ten model predictions. This reduces seed-dependent variance and makes the comparison less sensitive to a single lucky initialisation. We report both the ensemble and the distribution of individual seeds. The ensemble is the frozen model; the mean of single-seed metrics is not the same as the metric computed on the mean prediction.

Model fitting used ROC-AUC for early stopping, but the assessment of results was multi-metric. ROC-AUC measures the probability that a randomly drawn positive observation is ranked above a randomly drawn negative one. At 2.15 percent prevalence, ROC-AUC can conceal low precision, so PR-AUC is a central measure. In addition we report standardised partial ROC-AUC at a false-positive rate of up to 10 percent, Brier score, log-loss, calibration slope, calibration-in-the-large, expected calibration error and capture at fixed investigation budgets. The partial ROC-AUC is standardised so that, within the FPR region, 0.5 corresponds to chance ranking and 1.0 to perfect ranking; the raw partial area on FPR ≤ 0.10 would have a maximum of 0.10.

Probabilities were calibrated with a logistic Platt transformation estimated exclusively on the validation set. Calibration changes the level and spread of the probabilities but not the ranking when the transformation is monotone. It is therefore possible for ROC-AUC and PR-AUC to be good while the absolute probability is wrong. Precisely this happened in FY2024, and it is treated as a finding, not as a cosmetic detail. The validation set has two roles — early stopping for each ensemble member and estimation of the calibrator — and the roles share the same 765 outcomes, which can produce a mildly optimistic calibrator. Since the level gate failed in FY2024 regardless, this does not change the study's verdict, but a production calibrator should be estimated on data that did not also govern early stopping.

i = (1/10) Σs=4251 pi,s   and   logit(pi,cal) = αval + βval · logit(p̄i)

The canary feature was deterministic pseudorandom and was included in the frozen feature contract. That makes the audit trail reproducible, but also means it was not merely an external diagnostic: the trees could, in principle, split on it. A finally deployable model must be refitted and validated without the canary. The results in this article describe the actually frozen 215-input model and do not conceal this deviation.

The canary cannot be read off with gain-based feature importance: a continuous noise column always offers split points to a greedy tree learner and attains positive gain without carrying information. A canary check on trees must therefore use permutation importance or another label-breaking measure. The frozen run did not store such a reading, and the canary's contribution consequently cannot be cleared retrospectively (section 10).

6. Falsification and leakage audit

The term "leakage-free" is often used too loosely. A column can be free of the label itself and still carry a perfect fingerprint of future status. We therefore defined leakage as any information path through which the model gains access to events, registry states, source provenance or selection mechanisms that were not available for the business in question at the origin. The definition also covers the selection procedure itself: a field list chosen with knowledge of evaluation labels is leakage, even if every individual column is point-in-time clean in isolation.

The audit regime combines static contracts and empirical falsification tests. Static contracts reject forbidden sources, require a unique key per organisation number and accounting year, verify that the number of post-origin rows is zero, and hash SQL, code, feature order and artefacts. Empirical tests verify that duplicates do not change the result, that future rows can be appended without altering historical features, that source flows show no suspicious calendar growth, and that company-specific values lose power when shuffled within preserved source and time strata.

Figure 2 / Safety architecture

Six gates between raw public event and model feature

Adversarial control
The gates reduce different error types. No single control can prove the absence of all leakage; the combination makes the most important plausible mechanisms explicit and falsifiable.

The gates are not theoretical precautions; each answers an error class that actually occurred and was measured during development. The prohibition on contemporaneous role snapshots (gate 02) derives from an identity key that had been disambiguated against a later registry snapshot and thereby carried survival information into the person graph. The prohibition on source and coverage flags (gate 04) derives from a sentinel value and geographic indicators that in practice identified the collection channel of recovered, predominantly dissolved companies. The retention screen and the ban on bankruptcy-derived history derive from features that grew with the source's deletion window rather than with company risk. The stratified shuffling (gate 05) derives from a feature family whose control arms showed that the entire contribution lay in the availability pattern, not in the values. The requirement of frozen contracts on both sides of any comparison derives from one feature selection that had read the evaluation year's labels, and from one stored reference prediction with a different column contract than the valid comparator.

Point-in-time and duplicate gates

The capital, recency and person-network builders reported zero post-origin rows, full key coverage and unique organisation-number–FY keys. The recency and capital families passed the future-append test: if later source rows are appended, earlier company-years must remain bitwise unchanged. Both families also passed the logical duplicate tests. This is stronger than writing a date condition in SQL; the test challenges whether the condition actually protects the finished artefact.

Retention was checked by comparing source flows and feature distributions from FY2021 to FY2024. Symmetric growth above 1.5 times triggers suspicion of a calendar or retention signal. The newly selected families passed this screen. An auditor-network rate with weaker methodological grounding was expressly excluded, even though it could have improved a development number. Negative and rejected results are part of the model's evidence, because they show that access to a feature did not automatically lead to admission.

Negative control for the accounts block

The largest residual concern was that historical accounts values came from several retrieval layers. If "where the value came from" was correlated with later outcomes, the model could learn the source instead of the economics. In a pre-declared single-seed falsification, the entire 57-feature historical accounts and quality block was shuffled jointly between companies within strata preserving accounting year, industry and observability pattern. 709,772 rows were moved to a different company, without label reading.

The genuine block beat the control by +0.04143 ROC-AUC, +0.16523 PR-AUC and 839 additional hits in the top 1 percent. This is not a complete proof of causality, but it refutes a simple explanation in which the whole accounts contribution is due to source or coverage patterns. An independent panel reconciliation of 14 central accounts measures further showed roughly 99.4–99.9 percent agreement within a tolerance of max(NOK 1,000, 0.1 percent) on overlapping observations.

For feature families with low coverage, development additionally used control arms in which only the availability mask, shuffled values with the missingness pattern preserved, or placebo copies of existing fields were offered to the model; one family in which the mask alone reproduced the entire gain was rejected before freezing. The shuffling control in this section is the frozen, pre-declared variant of the same principle.

Exact replay and immutable receipts

The final run stored one proof receipt, not all intermediate boosters and prediction files. The code, feature order, freeze and input artefacts were bound with SHA-256. Seed 42 was then re-run under the same contract. The best iteration was 320 in both runs, the maximum absolute difference between the evaluation predictions was 0.0, and the prediction hash was identical.

7. Results: discrimination and incremental value

The frozen ten-seed ensemble reached ROC-AUC 0.8830 and average precision (AP) 0.4112 in FY2024. Standardised partial ROC-AUC at a false-positive rate of up to 10 percent was 0.7680. The Brier score after validation-based calibration was 0.0161, and log-loss was 0.0724. All figures were computed on the same evaluation population of 375,688 companies and 8,082 observed failures. The tables, the status banner and metodegrunnlag.json give full precision; running text is rounded to three or four decimals. For reference, an information-free model that always predicts the prevalence attains a Brier score of p(1 − p) = 0.021050; the low absolute Brier value chiefly reflects the low prevalence and must be read relative to it, not as free-standing evidence of good calibration.

AP is used because the outcome is rare. A random ranking has an expected precision near the prevalence — the share of companies with an observed outcome in the window, 8,082/375,688 = 0.021513. An AP of 0.4112 does not mean that all companies above a given threshold carry 41 percent risk; it is a ranking-wide measure summarising precision at different sensitivity levels. Absolute probability must be judged by the calibration measures in the next section.

MeasureReference, 198Candidate, 215Difference
ROC-AUC0.8792040.882950+0.003746
Average precision (AP)0.3980330.411187+0.013153
Standardised partial AUC, FPR ≤ 10%0.7616330.768019+0.006387
Brier score0.0163430.016147−0.000195
Hits in top 1%2,3332,352+19
Hits in top 5%4,4044,514+110
Hits in top 10%5,3355,417+82

Paired company bootstrap

To isolate the difference between candidate and reference, both models were evaluated on the same companies, and the same company sample was used for both models in each of 1,000 bootstrap draws. FY2024 has one row per company, so the company cluster and the row coincide in this cohort. The intervals are percentile intervals from the paired difference distribution. They are not adjusted for the fact that seven measures are reported simultaneously; they should be read as per-measure 95 percent intervals, not as simultaneous coverage for the whole set, and no single conclusion in the study rests on the weakest of several intervals. The intervals also condition on the two frozen models and express only sampling uncertainty in the evaluation cohort; uncertainty from refitting the models (other seeds or training samples) is not included, nor — as section 10 discusses — is optimism from the adaptive development process. The Brier improvement, invisible on the percentage-point axis of Figure 3, was +0.000195 (95 percent CI +0.000149 to +0.000243).

Figure 3 / Paired effect

The candidate improves most measures; top 1 percent crosses zero

95% company bootstrap
Forest plot in two panels for the difference between candidate and reference Panel A shows three ranking measures, panel B three sensitivity measures at fixed capacity, each with 95 percent intervals in percentage points. All intervals lie above zero except the top-one-percent sensitivity interval, which crosses zero. A · Ranking measures (percentage points) ROC-AUC95% CI +0.293 to +0.467+0.375 Average precision95% CI +0.990 to +1.629+1.315 Partial ROC-AUC95% CI +0.511 to +0.773+0.639 −0.250+0.50+1.00+1.50+2.00 B · Sensitivity at fixed capacity (percentage points of failures) Sensitivity, top 1%95% CI −0.210 to +0.691+0.246 Sensitivity, top 5%95% CI +0.836 to +1.793+1.361 Sensitivity, top 10%95% CI +0.582 to +1.457+1.015 −0.250+0.50+1.00+1.50+2.00 Change in percentage points; positive values favour the candidate. Both panels share one scale.
Panel A shows ranking measures; panel B shows sensitivity at fixed capacity, in percentage points of the 8,082 observed failures. The measure types are not directly comparable as effect sizes across panels. The Brier improvement is too small to display on this axis and is reported in the text with its own interval. The square marks the top-1-percent measure, whose 95 percent interval [−0.210, +0.691] crosses zero; it is therefore wrong to claim a statistically established improvement at exactly this capacity. That precisely this measure crosses zero is expected rather than anomalous: a threshold metric at a hard cut-off is decided by a few discordant companies near the boundary (the difference is 19 hits among 3,756 selected companies), while ROC-AUC and AP integrate over all thresholds and are estimated with far smaller bootstrap variance. Two related quantities are reported for this measure: the point difference is +0.235 percentage points (19 of 8,082 failures), while the central estimate of the paired bootstrap distribution is +0.246 percentage points. For the other six measures the two coincide exactly; the discrepancy here arises because the threshold metric is recomputed within each bootstrap draw.

Risk concentration at fixed investigation capacity

A risk model is often used as a prioritisation tool: how many real events can be examined when the organisation only has capacity to review a limited share of the company population? The candidate captured 2,352 failures among the 3,756 highest-ranked companies. That corresponds to 62.62 percent precision and 29.10 percent sensitivity; precision was thus 29.11 times the prevalence (lift 29.11). At five percent capacity, 4,514 failures were captured; at ten percent, 5,417.

Figure 4 / Operational concentration

Share of observed failures captured at three review budgets

Whole companies
Sensitivity at top one, five and ten percent capacity Candidate and reference lie close at top one percent. The candidate has higher sensitivity at five and ten percent. The three budgets are discrete categories; no curve is drawn between them. 0%20%40%60%70% Sensitivity / share of 8,082 failures Top 1%Top 5%Top 10% 2,333 / 2,352 hits4,404 / 4,514 hits5,335 / 5,417 hits Reference 198Candidate 215
The point estimates show practical concentration, not an automatic decision rule. The three capacities are discrete, categorical budgets; no continuous capacity–capture curve is measured or implied between them. The top-1-percent difference between the models is not statistically established; the improvements at top 5 and 10 percent are supported by the paired intervals.

Stability across seeds

The individual models had a mean ROC-AUC of 0.8781 with a standard deviation of 0.0012, and a mean AP of 0.3934 with a standard deviation of 0.0030. The ensemble was better than every single member on both ranking measures. This is expected when averaging reduces idiosyncratic variation between the trees, but it also shows why the ensemble metric must not be confused with the mean of the seed metrics.

Figure 5 / Stability

Ten individual seeds and the frozen ensemble

Seeds 42–51
Distribution of ROC-AUC and average precision for ten seeds Ten points show the individual models. A diamond to the right shows that the ensemble has higher ROC-AUC and average precision than every individual model. ROC-AUCmean 0.878133 · SD 0.001172 0.8750.8770.8790.8810.883 ensemble 0.882950 Average precisionmean 0.393399 · SD 0.002982 0.3860.3920.3980.4040.410 ensemble 0.411187 single seedseed meanensemble
The ensemble is computed from the mean of the predictions and then evaluated. It can therefore lie above all single-seed metrics. Vertical placement of the points is jitter to separate overlapping seeds; the vertical position carries no statistical meaning.

8. Calibration, economic utility and the gate that failed

Discrimination answers who should be ranked higher. Calibration answers whether "three percent risk" actually means roughly three events per hundred comparable companies. That distinction is decisive in credit use. A poorly calibrated model can still prioritise reviews efficiently, but it should not be used directly as a pricing basis, loss estimate or absolute probability of default.

In FY2024 the observed event rate was 2.151 percent, while the mean calibrated prediction was 1.327 percent. The observed/expected ratio was thus 1.6216: the overall observed risk was 62 percent higher than the model expected. Calibration-in-the-large — the intercept of a logistic recalibration of the outcome on the predicted logit, reported on the logit scale — was +0.6616 and broke the locked bound |CITL| ≤ 0.25. At the same time, the calibration slope of 0.9638 lay within the 0.7–1.3 interval, and the ten-bin ECE of 0.00825 lay within the 0.01 bound. The model thus had reasonable relative spread but the wrong overall level.

Because the predictions are frozen and maturation until 31 December 2026 can only add observed events, the O/E ratio of 1.62 is a lower bound for the fully mature window: the underprediction of the overall level will persist or grow with maturation — it cannot mature away. The discrimination and AP figures, by contrast, can move in either direction.

Figure 6 / Calibration

Two calibration gates passed; the level gate failed

FY2024
Comparison of predicted and observed event level with three calibration gates Mean predicted risk was 1.327 percent against 2.151 percent observed. Calibration slope and ECE passed, while calibration-in-the-large failed. 0%1%2%2.5% 1.327%2.151% Mean predictedObserved Pre-declared gates Slope 0.9638Passed: 0.7–1.3 ECE10 0.00825Passed: ≤ 0.01 CITL +0.6616Failed: |CITL| ≤ 0.25
Low ECE at low prevalence can coexist with an important relative level error. O/E and calibration-in-the-large show that the model underpredicts the overall event level. It must be recalibrated on a mature, representative cohort before the score can be read as an absolute 24-month probability.

As a prioritisation model, the candidate concentrated outcomes better than indiscriminate prioritisation at thresholds from one to twenty percent in the stored analysis, but such analyses alone cannot fix a business rule. The cost of a false positive, the loss from a missed bankruptcy, case-handling capacity and explainability requirements must be defined by the concrete use. The practical conclusion is therefore twofold: the ranking can be tried in controlled shadow testing; the probability must be recalibrated and monitored before it enters pricing or credit limits.

9. Subgroups and where the model is weaker

Average performance can conceal segments in which the ranking is substantially worse. The age analysis shows the greatest weakness among companies younger than two years: ROC-AUC 0.781. For companies between two and five years, ROC-AUC was 0.877, and for companies between ten and twenty years, 0.906. AP also varies, but it is directly affected by differences in event prevalence between the groups — shown in the table — and cannot be compared without that context.

Company ageCompaniesFailuresEvent rateROC-AUCAP
Under 2 years54,6721,6913.09%0.7808670.288329
2 to under 5 years79,1142,3733.00%0.8768460.554010
5 to under 10 years88,4352,1492.43%0.8949290.420735
10 to under 20 years95,3751,3271.39%0.9064280.365802
20 years or more58,0925420.93%0.8950680.291086

Among industry divisions with at least 1,000 companies and 20 failures, ROC-AUC ranged from roughly 0.797 to 0.934. Neither the age nor the industry estimates carry their own bootstrap intervals in the frozen receipt; differences between neighbouring groups should therefore be read as descriptive heterogeneity signals, not as a ranking of which industries the model "masters". Before production use, minimum segment sizes, per-segment calibration and a fallback for high-uncertainty groups must be defined.

The youngest companies have both short accounts histories and fewer dated administrative events. The lower AUC is therefore mechanistically plausible: the most important data sources have simply had less time to observe the company. A product presenting a single risk score should surface this information limitation and not convey the same impression of certainty for a one-year-old company as for a ten-year-old one.

10. Limitations and remaining evidence

First, the FY2024 outcome is not fully mature. The administrative read-out stops at 24 August 2026, while a complete 24-month window after 31 December 2024 closes on 31 December 2026. Companies failing in the final months are provisionally treated as having no observed outcome. The headline results must therefore be updated when the window is complete. The discrimination and AP figures can move in either direction as late events are recorded; as section 8 explains, the O/E ratio of 1.62 is a lower bound and can only persist or grow.

Second, FY2024 is temporally held out but not untouched. The evaluation year was seen repeatedly in the broader development process, and — a distinct mechanism described in section 3 — the 24-month outcome window of the last development vintage overlaps the evaluation window, without a purge in the frozen run. The paired bootstrap is valid for the frozen candidate against the frozen reference on this cohort, but it cannot correct optimism from all earlier choices. The next strong piece of evidence is a pre-registered time cohort that is not opened until every design choice is locked, or an independent portfolio with the same origin and outcome definition.

Third, the accounts clock is secured by the FY−1 rule and pre-origin approval traces, but the frozen evidence chain does not hold document-level filing, publication and revision timestamps for every individual historical value. Source-adjusted shuffling and high agreement with an independent panel reduce the suspicion that the entire signal is provenance, but they do not replace full document-bound point-in-time provenance. Our claim is therefore limited to leakage resistance under the implemented tests.

Fourth, the proof run was produced with an older label SQL than the later, chronologically corrected version. A read-only comparison found eight fewer qualifying FY2022 rows, of which six were failures, and four fewer FY2024 rows, all failures, under the corrected chronology. No shared row changed event status. The difference is small relative to 375,688 evaluations, but a strict final reproduction must bind the model to the corrected label code and re-run the frozen analysis.

Fifth, the deterministic canary remained a model input. Its final contribution was not stored, and it therefore cannot be declared irrelevant in hindsight — and, as section 5 explains, a gain-based importance reading could not have cleared it even if it had been stored. The deployable architecture must be refitted without the canary and show that the performance and the bootstrap conclusions stand. Likewise, the person network carries an open question about the completeness of negative edges: a positive event table can prove observed connections, but not always that the absence of a connection means genuine absence.

Sixth, the frozen deliverable is a metrics-and-architecture receipt, not a complete production package. Intermediate boosters, prediction vectors, the training matrix and the bootstrap draws were deliberately not stored, to limit artefact growth. That was right for a development audit, but it means a deployable model file, dependency lock, scoring contract, monitoring regime and operational test must be built separately. Reliability-bin data were not stored either, so a full reliability curve cannot be drawn from the final receipt without a newly bound prediction artefact.

Seventh, absolute calibration is insufficient and performance is heterogeneous. An AUC of 0.781 among the youngest companies may be too weak for certain decisions. Production use therefore requires purpose-specific threshold selection, segment analysis, recalibration, monitoring of population and prevalence drift, and a clear human override process. The model is not validated for other legal forms, other countries, all types of company termination, or regulatory capital purposes.

Finally, no internal study can establish that the model "beats banks" or named credit bureaus without identical companies, origin dates, outcomes, maturity and measurement criteria. Such comparisons require a shared blind test. Our measurable claim is narrower and stronger: in the defined FY2024 cohort, the frozen candidate ranked the observed failures well and beat its frozen reference on the pre-reported headline measures.

ClaimEvidenceVerdict
Strong ranking in FY2024AUC 0.882950; AP 0.411187; top-5 sensitivity 55.85%Supported in the observed cohort
Better than the frozen referencePaired intervals above zero for AUC, AP, partial AUC, Brier and top 5/10%Supported, conditional on the comparison
Top 1% is significantly betterΔ sensitivity +0.246 pp (point difference +0.235 pp); CI −0.210 to +0.691Not established
Leakage-free in an absolute senseExtensive tests, but residual accounts and label provenance and a shared outcome windowCannot be proven absolutely
Finished 24-month probabilityO/E 1.6216 (a lower bound while the window matures); CITL gate failed; the cohort is immatureNot supported
Production-ready model packageNo deployable booster or complete scoring package in the receiptNot yet

11. Conclusion

Byggsikt Risiko 24 shows that self-curated public Norwegian registry, accounts and gazette data can deliver strong, practically useful ranking of bankruptcy and forced dissolution. On 375,688 FY2024 companies with 8,082 observed failures, the ten-seed ensemble reached ROC-AUC 0.8830 and AP 0.4112. The five percent highest-ranked companies contained 4,514 of the failures. The improvement over the frozen 198-feature reference was positive with paired 95 percent intervals for AUC, AP, partial AUC, Brier and sensitivity at five and ten percent capacity.

The most important scientific value, however, is not one decimal number. It lies in the fact that several attractive results were withdrawn when they proved to be driven by cohort selection, the wrong clock, source provenance, calendar variables or missingness patterns. The chosen model is what remains after time gates, source prohibitions, de-duplication, fit-only imputation, negative shuffling controls, hash binding and bit-exact replay.

The verdict is precise: the model is a strong, reproducible development candidate for risk ranking and controlled shadow testing. It is not yet a fully calibrated probability model or an autonomous credit decision. Before such use, the FY2024 window must mature, the corrected label code must be bound, the canary removed, the probability recalibrated, and a deployable package validated on a new locked cohort. Reporting these boundaries does not weaken the model. It is what makes the result defensible.

Citation, version and declarations

How to cite this study

Byggsikt Data. "Point-in-time prediction of bankruptcy and forced dissolution in Norwegian limited companies: development, adversarial audit and out-of-time validation of a 24-month risk model." Byggsikt, 29 August 2026. Preliminary validation study, version 1.0, English edition. https://www.byggsikt.no/rad/konkursrisiko-norske-smaforetak/en

@techreport{byggsikt2026risiko24en,
  author      = {{Byggsikt Data}},
  title       = {Point-in-time prediction of bankruptcy and forced dissolution in Norwegian limited companies},
  institution = {Byggsikt},
  year        = {2026},
  month       = {8},
  type        = {Preliminary validation study, version 1.0, English edition},
  url         = {https://www.byggsikt.no/rad/konkursrisiko-norske-smaforetak/en},
  note        = {English edition of the Norwegian original at https://www.byggsikt.no/rad/konkursrisiko-norske-smaforetak/. Proof receipt SHA-256 333a0762d10ba574...; evidence in metodegrunnlag.json}
}

Version and status

Version 1.0, published 29 August 2026. This page is the English edition of the Norwegian original, which is the version of record; the two editions carry identical numbers. The study is a preliminary validation study: the FY2024 outcome is right-censored until 31 December 2026, and the headline figures will be updated in a version 1.1 when the window is complete. All headline results and comparison figures are reproduced in metodegrunnlag.json. Certain process and falsification figures — among them the shuffling control and panel reconciliation in section 6, the architecture parameters in section 5, the subgroup range in section 9 and the label-SQL comparison in section 10 — exist so far only in the SHA-256-bound proof receipt and are not independently verifiable from the public evidence file; they are scheduled for inclusion in its next version. Changes occur only through new, dated versions of this page.

Conflicts of interest and funding

The study was conducted and funded by Byggsikt, which develops the model it describes, for use in its own services. There is an obvious self-interest in a positive result. The counterweight is methodological and explicit: pre-declared gates that were allowed to fail and are reported as failed, paired confidence intervals also where they weaken the conclusion, negative controls, and a verdict that limits use to ranking and shadow testing. No external clients have influenced the design or the reporting.

Data availability

The sanitised evidence file is public in metodegrunnlag.json and bound to the publication with cryptographic fingerprints. The raw data derive from publicly available Norwegian registry sources; Byggsikt's collection and interpretation infrastructure is proprietary and is not published. External full reproduction is therefore limited, and this is stated as an explicit limitation in section 10.

References and methodological sources

  1. Brønnøysund Register Centre. About the Register of Company Accounts. Public availability, filing process and ordinary deadline for annual accounts.
  2. Brønnøysund Register Centre. Announcements from the Bankruptcy Register. Content and publication of bankruptcy and forced-dissolution announcements.
  3. Brønnøysund Register Centre. Central Coordinating Register open-data documentation. Technical services, data fields and public licence.
  4. Collins GS, Reitsma JB, Altman DG, Moons KGM. The TRIPOD statement. BMJ. 2015;350:g7594.
  5. Moons KGM et al. PROBAST+AI. Updated tool for quality, risk of bias and applicability in prediction models.
  6. Saito T, Rehmsmeier M. The Precision-Recall Plot Is More Informative than the ROC Plot When Evaluating Binary Classifiers on Imbalanced Datasets. PLOS ONE. 2015;10(3):e0118432.
  7. Van Calster B et al. Calibration: the Achilles heel of predictive analytics. BMC Medicine. 2019;17:230.
  8. Efron B, Tibshirani RJ. An Introduction to the Bootstrap. Chapman & Hall/CRC; 1993.
  9. Friedman JH. Greedy Function Approximation: A Gradient Boosting Machine. Annals of Statistics. 2001;29(5):1189–1232.

The headline results in this publication are bound to the sanitised public evidence file; the process figures cited in sections 5, 6, 9 and 10 are bound to the proof receipt. Source code, raw data and certain collection procedures are not published because they form part of Byggsikt's proprietary data infrastructure. This limits external full reproduction and is therefore stated as an explicit limitation.