Valutus.

Article · First published February 2021

Judge a metric by what it causes.

ESG ratings are flawed. The usual conclusion is that they therefore do harm. That does not follow.

A measure has to be judged on what it sets in motion, not only on what it captures.

Written in February 2021, when S&P began publishing ESG scores publicly, and kept here because the argument did not depend on the occasion. The ratings have changed since. The question of how to judge one has not.

01Three lenses

Is it accurate? Is it improving? What does it set in motion?

01

Accuracy

The ratings miss enormous things

Existing metrics are wanting in many areas, and context-based frameworks are a real improvement on them. But sustainability metrics as currently built miss something bigger than any of that: the effect of actually using what a company provides.

02

Evolution

Accounting took centuries

Sustainability measurement has moved a long way since the Dow Jones Sustainability Index launched. Financial accounting took hundreds of years to reach its current and still imperfect state. Expecting sustainability accounting to arrive finished is not a fair test.

03

Catalytic effect

What the measure causes

The serious objection is not that ESG ratings are inaccurate. It is that they actively obscure real sustainability by standing in its place. That is plausible, and it is the argument worth taking apart, because it is the one that would actually matter if it held.

02What the ratings miss

The largest omission in conventional sustainability measurement is not an emissions boundary or a supply chain tier. It is the consequence of the product itself, in use.

If an online network is used to plan the overthrow of democratically elected governments, it does not matter what powers the data centers. Clean energy does not wash away dirty deeds.

What gets counted

Operations. Energy, water, waste, the footprint of running the company.

What rarely does

Use. What the product does once it is in the world, which for many companies is the largest effect they have in either direction.

This is an argument for better measurement, not for less of it. A rating that misses the biggest term is incomplete. Incomplete is not the same as worthless.

03The catalytic case

Catalytic impact is not one thing. At least three factors drive it, and wider availability of ESG scores pushes on all three.

a

Utility

Wider availability makes the scores easier to use, and ease is most of what determines whether a measure actually gets used to direct investment.

b

Impetus

More availability leads to more use, as does any connection between the scores and stock performance. That raises the incentive to research and argue about them, which raises use again.

c

Attention

This is where the objection lands: attention paid to ESG scores is attention taken from deeper measures. There is something to it. It is not zero-sum.

Most people, and most investors, pay no attention to sustainability metrics of any kind. Not ESG, not context-based, not anything. So when some form of sustainability measurement reaches mainstream awareness, it raises the salience of the whole subject. It grows the pie rather than dividing it.

Within that larger pie, attention to ESG scores may well come at the expense of deeper measures. Both can be true at once, and the net can still be positive. In the years around this piece, attention to conventional ESG scores rose, and so did interest in the measurement and valuation work that sits well beyond it.

04Does acting feel like enough?

The second objection is moral licensing: acting on an ESG score makes people feel they have done their bit, which reduces what they do next. There is evidence for this, particularly where the action is taken publicly.

There is also evidence running the other way, and it is older than the debate.

Self-perception theory

Daryl Bem's account, published in 1972, holds that people infer their own attitudes by observing their own behavior, much as an outside observer would.

Applied to environmental action, the implication is direct. Someone who shuts off an idling engine can come to see themselves as a person who cares about this, and behaves accordingly afterward. Doing the small thing changes the self-description, and the self-description drives the next thing.

Both effects are real. Which one dominates is an empirical question that varies by circumstance, and neither side of it is settled by asserting that the metric is imperfect.

05Where this lands

On balance, wider availability of ESG ratings and the attention that comes with it should be a positive for sustainability, in spite of the flaws in the ratings. Not because they are right. Because they promote movement in the right direction.

What we measure affects what we do, and if we measure the wrong thing, we will do the wrong thing.

Joseph Stiglitz, Nobel laureate in economicsMismeasuring Our Lives, 2010

The quotation cuts both ways, and it should. A bad measure does real damage. But the second half of the sentence is a reason to improve the measure and keep measuring, not a reason to wait for a perfect one before starting. More attention behind measuring environmental and social performance, even imperfectly, does more to improve it than the alternative on offer, which is measuring nothing.

Imperfect is not the same as useless.