In April 2024, Nature published one of the most-cited climate economics papers of the decade. In December 2025, its own authors retracted it.
By the time they did, the paper’s numbers had already travelled into filed board climate disclosures across the financial sector, by a route no single board chose and few have traced. Following that route exposes a kind of model risk most governance frameworks do not look for.
A paper, a scenario set, a disclosure
The paper was Kotz, Levermann and Wenz, “The economic commitment of climate change”. It projected steep global income losses from warming already locked in, drew hundreds of thousands of downloads and was covered widely in the press. Then post-publication review, prompted by Thomas Bearpark, Dylan Hogan, Solomon Hsiang and Christof Schötz, found the results turned on the economic data for a single country. Correcting Uzbekistan’s subnational figures for 1995 to 1999, controlling for the data-source transitions they contained and accounting for spatial autocorrelation widened the mid-century damage range from 11 to 29 per cent out to 6 to 31 per cent, and cut the probability of damages diverging across emission scenarios by 2050 from 99 to 90 per cent. The authors judged the change too large for a correction and withdrew the paper on 3 December 2025.
The trouble is where the original estimates had gone in the meantime. The Network for Greening the Financial System had drawn on the paper for the physical-risk estimates in its Phase V scenarios. Banks, insurers and large corporates run those scenarios through their own models to produce the physical-risk numbers in their TCFD and UK SRS disclosures, and some use them to calibrate risk appetite. Norges Bank Investment Management, in a discussion note dated 27 February 2026, rebuilt the analysis on its own top-down estimates while otherwise following the NGFS approach. Under a surprise-warming scenario it put the hit to equity values at 12 per cent, against 14 to 21 per cent on the originally published estimates, and flagged high uncertainty across model specifications. The direction of the original finding survived; the size of it did not.
Nobody in the chain built the flawed model
Look at that chain again. One academic paper fed a central-bank scenario set. The scenario set fed a population of corporate disclosures. Not one of the firms at the end of the line wrote the model whose estimates they were reporting. Each party relied, reasonably, on the layer above: the firm trusted the NGFS, the NGFS trusted the peer-reviewed literature, the literature was in Nature. Every link in that chain is defensible on its own. The chain as a whole carried a retracted result into signed accounts.
This is inherited model risk, exposure sitting in evidence the firm did not produce, has not validated and in most cases has never named. It survives a clean validation report, because model risk management, whether under the PRA’s SS1/23 or the older SR 11-7 tradition, points inward. It governs the models a firm owns. The provenance of an external scenario, or the source study three steps upstream, tends to fall outside the perimeter that anyone is formally accountable for.
The NGFS, to its credit, had flagged the Phase V physical-risk methodology as under active review. The warning existed. It simply never travelled far enough down the chain to reach the audit committee signing the disclosure. A caveat only governs if it reaches the people accountable for the disclosure, and this one stopped at the modeller.
The reassuring number is the one to distrust
The same source that documents the retraction contains a second study making the point from the opposite direction.
Economists at the European Central Bank studied four major European floods between 2021 and 2024, from the catastrophic July 2021 events in Germany, Belgium and the Netherlands to the Valencia flash floods of October 2024. Their method was unusually clean. They combined metre-level Copernicus satellite flood maps with AnaCredit, the ECB’s loan-level credit register, and compared firms just inside a flood boundary with firms just a few hundred metres outside it, so that being flooded was effectively random. On the surface, lending looked stable. Loans to affected firms rose by 3.5 to 5 per cent in the flood quarter as they drew on credit lines, then fell back by roughly the same amount, leaving no meaningful net change. A board reading the aggregate would have seen calm.
Underneath, credit quality deteriorated and stayed that way. Default rates on pre-existing loans to flooded firms rose by around 0.7 percentage points cumulatively over two quarters, which annualised is close to a doubling of the roughly 1.4 per cent baseline default rate among those firms. The aggregate reassured while the credit book quietly worsened.
The two studies point the same way from opposite ends. Whether it is an inherited scenario estimate carrying the authority of a central-bank network, or a loan book whose aggregate looks untroubled, the comforting number is the one an adversarial reviewer should reach for first. It is usually the number nobody has looked underneath.
Make provenance a governed control
None of this argues for abandoning external scenarios. No firm can build the NGFS from scratch, and inherited evidence is unavoidable and usually sensible. The failure is in treating the inheritance as settled once it arrives.
Three things move provenance from a modelling detail to a governed control. The first is an evidence register for material disclosures: a plain record of which external studies, scenarios and vendor datasets each disclosure actually rests on. Most firms cannot produce this today, which is itself the finding. The second is a set of re-validation triggers with a named owner, so that a retraction, a revision, a superseding vintage or a supervisory caveat sets off an assessment rather than passing unnoticed. If the Kotz retraction would not reach your audit committee inside a reporting cycle, that is the gap. The third is a mandate to challenge inherited evidence, held by someone whose standing does not depend on the models under review, distinct from the second-line function that validates the models the firm built itself.
Owen Vallis has spent two decades on the receiving end of this chain. Inside FCA and PRA regulated institutions he built and validated IFRS 9 expected credit loss frameworks, with a documented mechanism translating macro and physical-risk scenarios into probability-of-default adjustments. He designed asset-level multi-hazard physical risk modelling across 10,000 properties in the UK, Europe and the Middle East, worked directly with the NGFS Phase V scenarios and authored TCFD climate-risk disclosures at group level. The provenance question is one he has answered from the inside, as the person taking an external scenario set, running it through a bank’s own models and writing the result into the disclosures a board signs.
Every board assumes it would catch a retracted study before it reached the accounts. The firms whose disclosures rest on Phase V assumed the same. The real question is narrower and more uncomfortable. If one of the external sources your climate disclosures depend on were withdrawn tomorrow, would anyone in your governance chain know in time to act, and who owns that?
Diagnostic SGaaS
Find out what your disclosures are resting on
A Diagnostic SGaaS engagement maps where your disclosures rest on inherited evidence no one has re-validated, and installs the register, triggers and challenge mandate that close the gap, so a Provision 29 declaration or a UK SRS sign-off rests on evidence the board can stand behind. That is the conversation worth having while the next retraction is still someone else's.
Owen Vallis is the founder of Marentis Labs, the firm that originated Strategic Governance as a Service. He spent ten years as UK Head of Fiduciary Risk Management at Credit Suisse and holds active board roles in the charity and education sectors. Schedule a confidential discussion.