A four-step framework for improving insurance pricing fairness

Research finds standard fairness tests can conceal discrimination in insurance pricing, and calls for corrected audits across the pricing process

Insurance companies using standard statistical formulas to test their pricing models for discrimination can produce results that mask bias rather than expose it, according to new research presented at a joint workshop hosted by the Wharton School and UNSW Business School.

Insurers are increasingly required to prove their pricing algorithms do not unfairly disadvantage customers based on factors such as race or postcode. The research, conducted by Fei Huang, an Associate Professor at UNSW’s School of Risk and Actuarial Studies, and Giles Hooker, a Professor in Statistics and Data Science at the Wharton School of the University of Pennsylvania, shows that the maths commonly used to run that check is flawed and can give companies a false pass even when bias is present in their pricing.

Associate Professor Fei Huang presenting at the Responsible AI _ Analytics for Insurance Workshop.jpg
Associate Professor Fei Huang presenting at the Responsible AI & Analytics for Insurance Workshop, hosted by the Wharton School and UNSW Business School. Photo: Supplied

A/Prof. Huang told the Responsible AI & Analytics for Insurance Workshop that a correction to standard error calculations completely reversed fairness test outcomes in an empirical analysis of a sample of insurers. Applying the uncorrected, commonly used formula, found no evidence of proxy discrimination at any of the 34 companies tested using Illinois postcode level premium data. Once the formula was corrected, every single company showed a statistically significant proxy effect.

A/Prof. Huang noted the empirical example draws on a public ProPublica dataset with limitations, including its reliance on neighbourhood-level averages and inferred rather than directly observed demographic data. She stressed the analysis is meant to illustrate how the corrected method changes audit outcomes, not to draw conclusions about any individual insurer.

“In this empirical example, if you use this standard formula, no company is flagged, so every company has no proxy discrimination,” she told the audience. “But if you correct the formula, then all the companies are significant, everyone has a proxy impact, so that means actually the correction has a very significant impact on the results.”

Applying a further threshold test, 16 of the 34 companies met both fairness criteria under the corrected method, compared with none under the traditional approach.


Why standard fairness tests fail

The problem lies in an assumption baked into most statistical significance testing: that pricing errors are independent and randomly distributed. A/Prof. Huang said this assumption does not hold for insurance pricing, because premiums for a given customer profile are set deterministically rather than drawn at random from a population.

“Because the pricing outputs are deterministic, that means the standard error that you calculate shouldn’t be based on the traditional one, which assumes the errors are IID (Independent and Identically Distributed).” She said the gap between the corrected and uncorrected calculations was substantial, with the corrected standard error averaging 92% smaller than the naive estimate assuming independent samples.

A/Prof. Huang’s research responds directly to draft fairness testing regulations proposed by Colorado and New York State, both of which require insurers to run quantitative tests demonstrating that the use of external consumer data and information sources does not serve as a proxy for any protected class, potentially resulting in unfair or unlawful discrimination. A/Prof. Huang said neither regulator’s draft set out the fairness reasoning behind the tests in detail, leaving room for insurers to interpret the intent themselves.

The missing piece: what fairness actually means

A/Prof. Huang said a further gap in the draft rules was that neither specifies a definition of fairness, even though fairness is what the tests are ultimately meant to establish. She argued insurers and regulators need to agree on a specific fairness standard before running any statistical test, rather than treating the regression as a self-contained answer.

"Across the regulations, a clear fairness standard isn't usually specified," she said. "They set out the regressions, but I'd recommend we attach a fairness notion to them, or explain the rationale behind the tests." She also proposed flipping the underlying hypothesis structure of these tests, so that insurers must actively demonstrate fairness rather than merely fail to disprove it.

Learn more: How insurers can mitigate the discrimination risks posed by AI

Most tests start by assuming a model is fair unless proven otherwise, but failing to disprove fairness is not the same as confirming it. A/Prof. Huang said this is the tension at the heart of current testing, and argued the starting assumption should be flipped, so that insurers carry the burden of proving their pricing is fair, rather than simply surviving a test built to assume it. Furthermore, she suggested regulators should also set a tolerance threshold, moving away from binary pass-fail significance testing toward equivalence testing that accounts for sample size.

Machine learning accuracy trade-off larger than expected

Beyond testing methodology, A/Prof. Huang presented findings from structural economic modelling showing that small reductions in machine learning accuracy carry disproportionate financial consequences for both insurers and customers. Her research modelled the full pricing pipeline, from claims cost modelling through to demand modelling and final market premiums, arguing that fairness interventions applied only at the machine learning stage fail to guarantee fairness in the price a customer actually pays.

“What we found is that even a sub-1% drop in prediction accuracy can lower insurer profit by 3% to over 4%, and it doesn’t spread evenly across consumers; some gain while others lose, even when the average welfare effect looks small,” said A/Prof. Huang, who added that this reframes the conventional trade-off debate in the field. “We want to emphasise that the fairness and accuracy trade-off should be dropped,” she said. “Instead, we should look at the fairness and welfare trade-off to better understand the real impact on consumers, rather than just the accuracy of the machine learning algorithm, for example.”

"If you choose to be fair in terms of premiums, you can’t be fair in terms of markups, so you need to make a decision here"

FEI HUANG

A/Prof. Huang said her modelling also found that fairness cannot simply be pursued at the machine learning stage and assumed to flow through to the final premium, because demand modelling and price optimisation between the cost estimate and the market price introduce further deviation from risk-based pricing. She said this practice remains common in Australia, despite being banned or significantly restricted in several US states for more than a decade.

“The draft regulations identify the need for fairness testing but leave some room for insurers in interpreting the fairness rationale behind the proposed tests, creating an opportunity for research to provide greater conceptual clarity,” she explained.

Premium fairness and markup fairness cannot both be achieved

In other research, A/Prof. Huang said one of the clearest findings from her modelling was that insurers cannot simultaneously make the final premium and the underlying markup fair across groups, thereby forcing a choice that has largely gone unacknowledged in industry practice.

“Making prices equal across consumer groups removes price gaps, but can widen markup gaps,” said A/Prof. Huang, who noted that this creates a decision insurers and regulators need to confront directly rather than treat fairness as a single achievable outcome. “If you choose to be fair in terms of premiums, you can’t be fair in terms of markups, so you need to make a decision here,” she added. "That’s another difficulty."

Subscribe to BusinessThink for the latest research, analysis and insights from UNSW Business School

The same tension extended to broader regulatory bans on pricing practices, including bans on price optimisation across several US states and the United Kingdom’s more limited restriction on price walking, in which insurers raise renewal premiums for existing customers with unchanged risk profiles. A/Prof. Huang’s modelling found no consistent answer as to whether such bans help or hurt consumers overall, with outcomes depending heavily on the structure of the market in question.

A four-step framework for insurers

A/Prof. Huang used the workshop to launch the Fair Pricing Playbook, an open-source framework built around four sequential steps: defining which fairness criterion applies to a given regulatory and business context, building a pricing model designed to satisfy that criterion, measuring the resulting welfare impact on both consumers and the insurer, and auditing the deployed system using statistically corrected testing.

She said the framework and accompanying code have been made publicly available, along with underlying data, so insurers, researchers and regulators can apply and critique the approach. A/Prof. Huang said she hoped the release would prompt collaboration and further exploration rather than serve as a finished answer.

"AI does not change things, AI does not create a new problem, it just amplifies an existing one"

FEI HUANG

She concluded her presentation with an observation about artificial intelligence: it has not introduced a new fairness problem in insurance pricing but has amplified a long-standing one by increasing the scale and speed of pricing decisions. “AI does not change things, AI does not create a new problem, it just amplifies an existing one. It’s really the scale and the speed that make things different, but the fundamental issue is exactly the same,” she concluded.

5 recommendations for industry professionals

1. For insurance executives

  • Set an organisation-wide definition of fairness for each line of business before approving pricing tests.
  • Assign accountability for premium fairness, markup fairness and consumer welfare.
  • Require pricing governance to cover the full path from claims modelling to the market premium.

2. For actuaries and data teams

  • Review whether audit formulas assume independent and identically distributed errors.
  • Report effect size, confidence intervals and tolerance thresholds alongside significance results.
  • Test pricing outcomes after demand modelling and price optimisation, not only after claims modelling.

3. For risk and compliance teams

  • Map each regulatory requirement to a fairness criterion, method, threshold and evidence source.
  • Commission audits that challenge assumptions within the test, not only the model under review.
  • Retain documentation covering data, proxy variables, model decisions, welfare effects and deployment reviews.

4. For regulators and policymakers

  • Define what fairness means before prescribing regression tests.
  • State the statistical basis, hypothesis structure and threshold required for compliance.
  • Assess effects on premiums, markups, insurer profit and consumer welfare.

5. For boards

  • Ask management which fairness outcome the pricing system seeks to achieve.
  • Require evidence that the audit method suits pricing outputs.
  • Treat insurance pricing fairness as a governance issue involving customers, regulation and financial performance.

Republish

You are free to republish this article both online and in print. We ask that you follow some simple guidelines.

Please do not edit the piece, ensure that you attribute the author, their institute, and mention that the article was originally published on Business Think.

By copying the HTML below, you will be adhering to all our guidelines.

Press Ctrl-C to copy